Rule-Based Aspect Extraction and IndoBERT-Based Sentiment Classification of Ruparupa Mobile Application Reviews
DOI:
https://doi.org/10.63158/journalisi.v8i4.1816Keywords:
IndoBERT, Rule-Based Aspect Extraction, Weak Labeling, Google Play Review Mining, E-Commerce ReviewsAbstract
This study evaluates a pipeline that separates rule-based aspect extraction from IndoBERT-based binary sentiment classification for Indonesian Ruparupa mobile application reviews. Google Play reviews were collected on 31 July 2026, anonymized, deduplicated before aspect expansion, and cleaned by lowercasing, removing URLs, emails, and special characters, and normalizing whitespace; no slang normalization, stop-word removal, or stemming was applied. Ratings 1-2 and 4-5 provided weak negative and positive labels, while three-star reviews were excluded. A 331-entry aspect dictionary mapped 1,495 unique reviews into 2,873 aspect-review pairs across six aspects. Across five repeated leakage-free group hold-out splits, IndoBERT achieved mean accuracy 0.9179 ± 0.0214, macro F1 0.9178 ± 0.0214, and ROC-AUC 0.9719 ± 0.0103; a calibrated TF-IDF + linear SVM baseline achieved 0.8765 ± 0.0124, 0.8759 ± 0.0127, and 0.9452 ± 0.0102, respectively. A McNemar test on run 1 showed a significant paired difference (p = 0.00013). Performance measures agreement with rating-derived weak labels rather than human-validated aspect sentiment. Because results from system-assigned aspects lacked independent human validation, aspect frequencies are descriptive rule-system outputs. Within this dataset, IndoBERT performed consistently across the five splits; supervised aspect extraction and human aspect-level annotation remain priorities.
Downloads
References
[1] J. Dąbrowski, E. Letier, A. Perini, and A. Susi, “Analysing app reviews for software engineering: A systematic literature review,” Empir. Softw. Eng., vol. 27, art. no. 43, 2022, doi: 10.1007/s10664-021-10065-7.
[2] R. Massenon et al., “Mobile app review analysis for crowdsourcing of software requirements: A mapping study of automated and semi-automated tools,” PeerJ Comput. Sci., vol. 10, art. no. e2401, 2024, doi: 10.7717/peerj-cs.2401.
[3] B. Liu, Sentiment Analysis and Opinion Mining. in Synthesis Lectures on Human Language Technologies. Morgan & Claypool Publishers, 2012. doi: 10.2200/S00416ED1V01Y201204HLT016.
[4] W. Zhang, X. Li, Y. Deng, L. Bing, and W. Lam, “A Survey on Aspect-Based Sentiment Analysis: Tasks, Methods, and Challenges,” IEEE Trans. Knowl. Data Eng., vol. 35, no. 11, pp. 11019–11038, 2023, doi: 10.1109/TKDE.2022.3230975.
[5] M. Pontiki, D. Galanis, J. Pavlopoulos, H. Papageorgiou, I. Androutsopoulos, and S. Manandhar, “SemEval-2014 Task 4: Aspect Based Sentiment Analysis,” in Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), Dublin, Ireland: Association for Computational Linguistics, 2014, pp. 27–35. doi: 10.3115/v1/S14-2004.
[6] M. Pontiki et al., “SemEval-2016 Task 5: Aspect Based Sentiment Analysis,” in Proceedings of the 10th International Workshop on Semantic Evaluation (SemEval-2016), San Diego, California: Association for Computational Linguistics, 2016, pp. 19–30. doi: 10.18653/v1/S16-1002.
[7] Y. C. Hua, P. Denny, K. Taskova, and J. Wicker, “A Systematic Review of Aspect-Based Sentiment Analysis: Domains, Methods, and Trends,” Artif. Intell. Rev., vol. 57, no. 11, art. no. 296, 2024, doi: 10.1007/s10462-024-10906-z.
[8] D. R. I. M. Setiadi, W. Warto, A. R. Muslikh, K. Nugroho, and A. N. Safriandono, “Aspect-Based Sentiment Analysis on E-commerce Reviews using BiGRU and Bi-Directional Attention Flow,” J. Comput. Theor. Appl., vol. 2, no. 4, pp. 470–480, 2025, doi: 10.62411/jcta.12376.
[9] E. Yulianti and N. K. Nissa, “ABSA of Indonesian Customer Reviews Using IndoBERT: Single-Sentence and Sentence-Pair Classification Approaches,” Bull. Electr. Eng. Informatics, vol. 13, no. 5, pp. 3579–3589, 2024, doi: 10.11591/eei.v13i5.8032.
[10] S. Imron, E. I. Setiawan, J. Santoso, and M. H. Purnomo, “Aspect Based Sentiment Analysis Marketplace Product Reviews Using BERT, LSTM, and CNN,” J. RESTI (Rekayasa Sist. dan Teknol. Informasi), vol. 7, no. 3, pp. 586–591, 2023, doi: 10.29207/resti.v7i3.4751.
[11] M. T. A. Bangsa, S. Priyanta, and Y. Suyanto, “Aspect-Based Sentiment Analysis of Online Marketplace Reviews Using Convolutional Neural Network,” IJCCS (Indonesian J. Comput. Cybern. Syst.), vol. 14, no. 2, pp. 123–134, 2020, doi: 10.22146/ijccs.51646.
[12] A. Vaswani et al., “Attention is all you need,” in Advances in Neural Information Processing Systems, 2017, pp. 5998–6008.
[13] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “BERT: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, MN, USA, 2019, pp. 4171–4186. doi: 10.18653/v1/N19-1423.
[14] B. Wilie et al., “IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding,” in Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing, Suzhou, China, 2020, pp. 843–857.
[15] F. Koto, A. Rahimi, J. H. Lau, and T. Baldwin, “IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP,” in Proceedings of the 28th International Conference on Computational Linguistics, Barcelona, Spain, 2020, pp. 757–770. doi: 10.18653/v1/2020.coling-main.66.
[16] E. Mulyati, M. I. C. Rachmatullah, and A. S. Firmansyah, “Sentiment Analysis of Pospay Application Reviews Using the BERT Deep Learning Method,” J. Tek. Inform., vol. 18, no. 2, pp. 173–183, 2025, doi: 10.15408/jti.v18i2.41116.
[17] A. Ratner, S. H. Bach, H. Ehrenberg, J. Fries, S. Wu, and C. Ré, “Snorkel: Rapid Training Data Creation with Weak Supervision,” VLDB J., vol. 29, pp. 709–730, 2020, doi: 10.1007/s00778-019-00552-1.
[18] S. Kapoor and A. Narayanan, “Leakage and the Reproducibility Crisis in Machine-Learning-Based Science,” Patterns, vol. 4, no. 9, art. no. 100804, 2023, doi: 10.1016/j.patter.2023.100804.
[19] S. Kapoor et al., “REFORMS: Consensus-Based Recommendations for Machine-Learning-Based Science,” Sci. Adv., vol. 10, no. 18, art. no. eadk3452, 2024, doi: 10.1126/sciadv.adk3452.
[20] G. Salton and C. Buckley, “Term-Weighting Approaches in Automatic Text Retrieval,” Inf. Process. Manag., vol. 24, no. 5, pp. 513–523, 1988, doi: 10.1016/0306-4573(88)90021-0.
[21] C. Cortes and V. Vapnik, “Support-Vector Networks,” Mach. Learn., vol. 20, pp. 273–297, 1995, doi: 10.1007/BF00994018.
[22] F. Pedregosa et al., “Scikit-Learn: Machine Learning in Python,” J. Mach. Learn. Res., vol. 12, no. 85, pp. 2825–2830, 2011.
[23] M. Hoang, O. A. Bihorac, and J. Rouces, “Aspect-Based Sentiment Analysis Using BERT,” in Proceedings of the 22nd Nordic Conference on Computational Linguistics, Turku, Finland: Linköping University Electronic Press, 2019, pp. 187–196.
[24] H. N. Alfiana, A. Doewes, and B. Widoyono, “Aspect-Based Sentiment Analysis of Access by KAI Application Reviews Using IndoBERT for Multi-Label Classification Tasks,” J. Tek. Inform., vol. 7, no. 1, pp. 286–306, 2026, doi: 10.52436/1.jutif.2026.7.1.5402.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Information Systems and Informatics

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors Declaration
- The Authors certify that they have read, understood, and agreed to the Journal of Information Systems and Informatics (JournalISI) submission guidelines, policies, and submission declaration. The submission has been prepared using the provided template.
- The Authors certify that all authors have approved the publication of this manuscript and that there is no conflict of interest.
- The Authors confirm that the manuscript is their original work, has not received prior publication, is not under consideration for publication elsewhere, and has not been previously published.
- The Authors confirm that all authors listed on the title page have contributed significantly to the work, have read the manuscript, attest to the validity and legitimacy of the data and its interpretation, and agree to its submission.
- The Authors confirm that the manuscript is not copied from or plagiarized from any other published work.
- The Authors declare that the manuscript will not be submitted for publication in any other journal or magazine until a decision is made by the journal editors.
- If the manuscript is finally accepted for publication, the Authors confirm that they will either proceed with publication immediately or withdraw the manuscript in accordance with the journal’s withdrawal policies.
- The Authors agree that, upon publication of the manuscript in this journal, they transfer copyright or assign exclusive rights to the publisher, including commercial rights














