Predicting Higher Sales Opportunities for Consumer Products Compared to the Previous Year Using Apriori and Random Forest with SHAP Interpretation

Authors

  • Syaiful Rahman Lubis Universitas Pembangunan Panca Budi
  • Muhammad Syahputra Novelan Universitas Pembangunan Panca Budi

Keywords:

Apriori, Random Forest, SHAP, Sales Prediction, Interpretable Machine Learning

Abstract

This study aims to develop a predictive model for the likelihood of increased consumer product sales in the current year compared to the previous year by utilizing a combination of the Apriori and Random Forest algorithms, as well as model interpretation using SHAP (SHapley Additive exPlanations). The research methodology consists of data exploration, extraction of association patterns between products using the Apriori algorithm, feature engineering based on annual sales history, classification using Random Forest, and interpretation of feature contributions using SHAP. The data used consists of consumer product sales transaction data covering year, product name, and quantity sold. The results show that the Random Forest model is capable of classifying products with the potential for increased sales with good accuracy and ROC-AUC values. The Apriori analysis generated a number of association rules between products with significant support, confidence, and lift values, while the SHAP interpretation indicated that the previous year’s sales and transaction frequency are the most influential factors in predicting sales growth. This study is expected to assist business owners in developing inventory and product promotion strategies based on predictions that can be explained transparently.

References

P. H. Putra, D. Selvida, and M. S. Novelan, “Application of Apriori Algorithm in Data Mining to Find Consumer Purchasing Patterns in Supermarkets,” J. Komput. Teknol. Inf. Sist. Inf., vol. 4, no. 1, pp. 221–226, 2025, doi: 10.62712/juktisi.v4i1.392.

R. Hermawan and M. Iqbal, “JoCoSiR,” vol. 2337, pp. 1–9.

S. D. Putri and S. Sitohang, “Analisis Pola Pembelian Konsumen Menggunakan Algoritma Apriori,” Comput. Sci. Ind. Eng., vol. 9, no. 7, pp. 285–290, 2023, doi: 10.33884/comasiejournal.v9i7.7889.

M. J. Faisti, R. H. Kusumodestoni, and G. W. N. Wibowo, “Mental Health Classification Using Naïve Bayes and Random Forest Algorithms,” J. Appl. Informatics Comput., vol. 9, no. 4, pp. 1740–1750, 2025, doi: 10.30871/jaic.v9i4.10144.

S. I. Oktora, D. Matualage, K. A. Notodiputro, and B. Sartono, “Data-Driven Insights Into Underdeveloped Regencies: SHAP-Based Explainable Artificial Intelligence Approach,” Int. J. Artif. Intell. Res., vol. 9, no. 1, 2025, doi: 10.29099/ijair.v9i1.1399.

D. S. Nugroho, N. Islahudin, V. Normasari, and S. Z. Al Hakiim, “Penerapan Market Basket Analysis (Mba) Data Mining Menggunakan Metode Asosiasi Appriori Dan Fp-Growth Pada Wan Caffeine Addict Yogyakarta,” JISI J. Integr. Sist. Ind., vol. 11, no. 1, pp. 121–134, 2024, doi: 10.24853/jisi.11.1.121-134.

R. Wilda, D. Saripurna, and O. K. Sulaiman, “Implementasi Algoritma Frequent Pattern Growth (FP-Growth) untuk Pola Penjualalan Tiket Travel pada PT Taxi Kita Bersama,” Hello World J. Ilmu Komput., vol. 3, no. 3, pp. 138–145, 2025, doi: 10.56211/helloworld.v3i3.588.

D. Alfitra, M. Afdal, M. Fronita, and E. Saputra, “Analisa Keranjang Belanja untuk Menentukan Tata Letak Barang Menggunakan Algoritma FP-Growth,” J. Sist. Inf., vol. 13, no. 4, pp. 1651–1661, 2024, [Online]. Available: http://sistemasi.ftik.unisi.ac.id

Eko Budianto and Muhammad Iqbal, “Model Predictive Analysis of Performance in Training and Course Institutions Using Naive Bayes and K-Means Clustering,” J. Comput. Sci. Res., vol. 3, no. 1, pp. 10–16, 2025, doi: 10.65126/jocosir.v3i1.68.

N. D. Ariyanta, A. N. Handayani, J. T. Ardiansah, and K. Arai, “Ensemble learning approaches for predicting heart failure outcomes: A comparative analysis of feedforward neural networks, random forest, and XGBoost,” Appl. Eng. Technol., vol. 3, no. 3, pp. 173–184, 2024, doi: 10.31763/aet.v3i3.1750.

A. Ernawati, Z. Sitorus, M. Iqbal, and D. Nasution, “Penerapan Data Mining Untuk Klasifikasi Penduduk Miskin Di Kabupaten Labuhanbatu Menggunakan Random Forest Dan K-Nearest Neighbors,” Bull. Inf. Technol., vol. 6, no. 2, pp. 23–35, 2025, doi: 10.47065/bit.v5i2.1783.

M. S. Novelan, S. Efendi, P. Sihombing, and H. Mawengkang, “Vehicle Routing Problem Optimization With Machine Learning in Imbalanced Classification Vehicle Route Data,” Eastern-European J. Enterp. Technol., vol. 5, no. 3(125), pp. 49–56, 2023, doi: 10.15587/1729-4061.2023.288280.

J. Anggara, F. R. Ramadhan, A. F. Syofian, A. A. Panjaitan, and R. F. Fithra, “Pengembangan Sistem Prediksi Harga dan Rekomendasi Mobil Bekas Berbasis Machine Learning,” J. Technol. Informatics, vol. 7, no. 1, pp. 11–23, 2025, doi: 10.37802/joti.v7i1.987.

S. Darma, A. Jihad, A. Fayed, S. M. P. Pardede, and M. H. Aqsha, “Predictive Analysis of Flood Risk Factors Based on a Machine Learning Approach : Comparative Study of SVM and XGBoost Algorithms,” vol. 3, no. 1, pp. 24–33, 2026.

N. Alfarizi, P. Lydia, M. S. Novelan, A. Putra, and S. Sinurat, “Comparative Machine Learning Analysis for Sentiment Classification of Sumatra Disaster 2025,” vol. 3, no. 1, pp. 68–76, 2026.

Downloads

Published

2025-10-27

Most read articles by the same author(s)