Cross-Cohort Early Prediction of Cluster-Derived GPA Trajectories Using First-Year Academic Data and Machine Learning
DOI:
https://doi.org/10.63158/journalisi.v8i4.1871Keywords:
Academic risk, Cohort drift, Cross-cohort validation, K-Means trajectory labeling, Student success predictionAbstract
Early identification of unstable academic development can support timely advising, yet models evaluated only with random splits may not transfer across cohorts. This study examines whether data available after semester-2 grades are finalized can classify two K-Means-derived GPA trajectory labels—stable-improving and fluctuating—observed in semesters 3–7. A retrospective dataset of 1,362 anonymized students was analyzed. Labels were generated from standardized semester 3–7 GPA profiles, while predictors were limited to admission attributes and semester 1–2 records. Train-only preprocessing, RFECV, validation-set threshold tuning, calibration analysis, and GaussianNB, logistic regression, Random Forest, and SVM-RBF were evaluated using a stratified random test and a 2024/2025 administrative-cohort holdout. Random Forest achieved 0.831 balanced accuracy and 0.916 ROC-AUC on the random test. In the holdout, where fluctuating-label prevalence increased from 25.5% to 69.3%, Random Forest had the highest threshold-based balanced accuracy (0.581), whereas SVM-RBF had the highest ROC-AUC (0.763). Drift was substantial in student type, admission semester, status, and study-program composition. In this dataset, random splitting substantially overestimated cross-cohort performance. The observed degradation represents a moderate generalization challenge; these cluster-derived labels require cohort-aware recalibration and educational validation, and predictions should be used only as a human-supervised screening aid.
Downloads
References
[1] D. Ifenthaler and J. Yau, “Utilising learning analytics to support study success in higher education: a systematic review,” Educ. Technol. Res. Dev., vol. 68, pp. 1961–1990, 2020.
[2] A. Namoun and A. Alshanqiti, “Predicting Student Performance Using Data Mining and Learning Analytics Techniques: A Systematic Literature Review,” Appl. Sci., vol. 11, no. 1, p. 237, 2021.
[3] E. A. Alyahyan and D. Düştegör, “Predicting academic success in higher education: literature review and best practices,” Int. J. Educ. Technol. High. Educ., vol. 17, pp. 1–21, 2020.
[4] N. Sghir, A. Adadi, and M. Lahmer, “Recent advances in Predictive Learning Analytics: A decade systematic review (2012--2022),” Educ. Inf. Technol., vol. 28, pp. 8299–8333, 2023.
[5] Á. Kocsis and G. Molnár, “Factors influencing academic performance and dropout rates in higher education,” Oxford Rev. Educ., vol. 51, pp. 414–432, 2024.
[6] F. Marbouti, J. Ulas, and C. Wang, “Academic and Demographic Cluster Analysis of Engineering Student Success,” IEEE Trans. Educ., vol. 64, pp. 261–266, 2021.
[7] T. Baron et al., “Signatures of medical student applicants and academic success,” PLoS One, vol. 15, 2020.
[8] N. Nawa et al., “Associations between demographic factors and the academic trajectories of medical students in Japan,” PLoS One, vol. 15, no. 5, p. e0233371, 2020.
[9] E. Alhazmi and A. M. Sheneamer, “Early Predicting of Students Performance in Higher Education,” IEEE Access, vol. 11, pp. 27579–27589, 2023.
[10] A. F. Mohamed Nafuri, N. S. Sani, N. F. A. Zainudin, A. H. A. Rahman, and M. Aliff, “Clustering Analysis for Classifying Student Academic Performance in Higher Education,” Appl. Sci., vol. 12, no. 19, p. 9467, 2022.
[11] A. Ridwan, T. Sutikno, I. Riyadi, and W. C. Wahyudin, “On-Time Student Graduation Prediction Modeling: A Comparative Analysis of Naive Bayes Algorithm and Other Data Mining Classifications: Pemodelan Prediksi Kelulusan Mahasiswa Tepat Waktu: Analisis Komparatif Algoritma Naive Bayes Dan Klasifikasi Data Mining Lainnya,” JOINCS (Journal Informatics, Network, Comput. Sci., vol. 8, no. 2, pp. 128–135, 2025.
[12] C. Foster and P. Francis, “A systematic review on the deployment and effectiveness of data analytics in higher education to improve student outcomes,” Assess. Eval. High. Educ., vol. 45, pp. 822–841, 2020.
[13] L. Breiman, “Random Forests,” Mach. Learn., vol. 45, no. 1, pp. 5–32, 2001, doi: 10.1023/A:1010933404324.
[14] I. Guyon, J. Weston, S. Barnhill, and V. Vapnik, “Gene selection for cancer classification using support vector machines,” Mach. Learn., vol. 46, no. 1, pp. 389–422, 2002.
[15] C. Cortes and V. Vapnik, “Support-Vector Networks,” Mach. Learn., vol. 20, no. 3, pp. 273–297, 1995.
[16] T. Fawcett, “An introduction to ROC analysis,” Pattern Recognit. Lett., vol. 27, no. 8, pp. 861–874, 2006, doi: 10.1016/j.patrec.2005.10.010.
[17] T. Saito and M. Rehmsmeier, “The precision-recall plot is more informative than the ROC plot when evaluating binary classifiers on imbalanced datasets,” PLoS One, vol. 10, no. 3, p. e0118432, 2015.
[18] K. H. Brodersen, C. S. Ong, K. E. Stephan, and J. M. Buhmann, “The balanced accuracy and its posterior distribution,” in 20th International Conference on Pattern Recognition, 2010, pp. 3121–3124.
[19] G. Feng, M. Fan, and Y. Chen, “Analysis and Prediction of Students’ Academic Performance Based on Educational Data Mining,” IEEE Access, vol. 10, pp. 19558–19571, 2022.
[20] H. Lim, S. Kim, K.-M. Chung, K. Lee, T. Kim, and J. Heo, “Is college students’ trajectory associated with academic performance?,” Comput. Educ., vol. 178, p. 104397, 2022.
[21] K. M. L. Jones et al., “We’re being tracked at all times: Student perspectives of their privacy in relation to learning analytics in higher education,” J. Assoc. Inf. Sci. Technol., vol. 71, pp. 1044–1059, 2020.
[22] K. M. L. Jones et al., “Transparency and Consent: Student Perspectives on Educational Data Analytics Scenarios,” portal Libr. Acad., vol. 23, pp. 485–515, 2023.
[23] P. Prinsloo, S. Slade, and M. Khalil, “The answer is (not only) technological: Considering student data privacy in learning analytics,” Br. J. Educ. Technol., vol. 53, pp. 876–893, 2022.
[24] C. Fachola, A. Tornaría, P. Bermolen, G. Capdehourat, L. Etcheverry, and M. Fariello, “Federated Learning for Data Analytics in Education,” Data, vol. 8, p. 43, 2023.
[25] W. C. Wahyudin, T. Sutikno, and R. Umar, “A Cluster-Label-Based Framework for Water Quality Risk Pattern Classification Using Naive Bayes and Random Forest,” J. Ilm. Ilmu Terap. Univ. Jambi, vol. 10, no. 3, pp. 1511–1522, 2026, doi: 10.22437/jiituj.v10i3.57126.
[26] M. Yakubu and A. Abubakar, “Applying machine learning approach to predict students’ performance in higher educational institutions,” Kybernetes, vol. 51, no. 2, pp. 916–934, 2021, doi: 10.1108/K-12-2020-0865.
[27] M. V Martins, D. Tolledo, J. Machado, L. M. T. Baptista, and V. Realinho, “Early Prediction of Student’s Performance in Higher Education: A Case Study,” in Trends and Applications in Information Systems and Technologies, WorldCIST 2021, 2021, vol. 1365, pp. 166–175, doi: 10.1007/978-3-030-72657-7_16.
[28] M. Yağcı, “Educational data mining: prediction of students’ academic performance using machine learning algorithms,” Smart Learn. Environ., vol. 9, p. 11, 2022, doi: 10.1186/s40561-022-00192-z.
[29] Y. S. Balcıoğlu and M. Artar, “Predicting academic performance of students with machine learning,” Inf. Dev., vol. 41, no. 3, pp. 896–915, 2025, doi: 10.1177/02666669231213023.
[30] K. Mahawar and P. Rattan, “Empowering education: Harnessing ensemble machine learning approach and ACO-DT classifier for early student academic performance prediction,” Educ. Inf. Technol., vol. 30, pp. 4639–4667, 2025, doi: 10.1007/s10639-024-12976-6.
[31] N. Butt, Z. Mahmood, K. Shakeel, S. Alfarhood, M. S. Safran, and I. Ashraf, “Performance Prediction of Students in Higher Education Using Multi-Model Ensemble Approach,” IEEE Access, vol. 11, pp. 136091–136108, 2023, doi: 10.1109/ACCESS.2023.3336987.
[32] S. Alwarthan, N. Aslam, and I. U. Khan, “Predicting Student Academic Performance at Higher Education Using Data Mining: A Systematic Review,” Appl. Comput. Intell. Soft Comput., vol. 2022, p. 8924028, 2022, doi: 10.1155/2022/8924028.
[33] E. Ahmed, “Student Performance Prediction Using Machine Learning Algorithms,” Appl. Comput. Intell. Soft Comput., vol. 2024, p. 4067721, 2024, doi: 10.1155/2024/4067721.
[34] A. Harif and M. A. Kassimi, “Predictive Modeling of Student Performance Using RFECV-RF for Feature Selection and Machine Learning Techniques,” Int. J. Adv. Comput. Sci. Appl., vol. 15, no. 7, pp. 228–237, 2024, doi: 10.14569/IJACSA.2024.0150723.
[35] S. Batool, J. Rashid, M. W. Nisar, J. Kim, H.-Y. Kwon, and A. Hussain, “Educational data mining to predict students’ academic performance: A survey study,” Educ. Inf. Technol., vol. 28, pp. 905–971, 2023, doi: 10.1007/s10639-022-11152-y.
[36] E. Tiukhova et al., “Explainable Learning Analytics: Assessing the stability of student success prediction models by means of explainable AI,” Decis. Support Syst., vol. 182, p. 114229, 2024, doi: 10.1016/j.dss.2024.114229.
[37] R. Alamri and B. Alharbi, “Explainable Student Performance Prediction Models: A Systematic Review,” IEEE Access, vol. 9, pp. 33132–33143, 2021, doi: 10.1109/ACCESS.2021.3061368.
[38] F. Arévalo-Cordovilla and M. Peña, “Evaluating ensemble models for fair and interpretable prediction in higher education using multimodal data,” Sci. Rep., vol. 15, art. 29420, 2025, doi: 10.1038/s41598-025-15388-9.
[39] W. C. Wahyudin, T. Sutikno, R. Umar, and A. Ridwan, “Comparison of Data Mining Model Performance in Heart Disease Detection with Feature Selection Application,” JOINCS (Journal Informatics, Network, Comput. Sci., vol. 8, no. 1, pp. 87–93, 2025, doi: 10.21070/joincs.v8i1.1669.
[40] W. C. Wahyudin, T. Sutikno, and R. Umar, “Identification of Bengawan Solo River Water Quality Patterns Using K-Means Clustering Based on Physicochemical and Environmental Parameters,” JOINCS (Journal Informatics, Network, Comput. Sci., vol. 9, no. 1, pp. 43–48, 2026, doi: 10.21070/joincs.v9i1.1710.
[41] L. N. Hakim, F. M. Hana, and W. C. Wahyudin, “Klasifikasi Komentar Toksik Berbahasa Indonesia di Media Sosial Berbasis Fine-Tuning IndoBERT,” JURIKOM (Jurnal Ris. Komputer), vol. 13, no. 1, pp. 202–209, 2026, doi: 10.30865/jurikom.v13i1.9449.
[42] S. P. Afrisia, F. M. Hana, and W. C. Wahyudin, “Implementasi Metode Long Short Term Memory (LSTM) pada Chatbot Kesehatan Mental Mahasiswa,” Sainteks, vol. 21, no. 2, pp. 107–116, 2024, doi: 10.30595/sainteks.v21i2.23869.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Journal of Information Systems and Informatics

This work is licensed under a Creative Commons Attribution 4.0 International License.
Author's Declaration
- The Authors certify that they have read, understood, and agreed to the Journal of Information Systems and Informatics (JournalISI) submission guidelines, policies, and submission declaration. The submission has been prepared using the provided template.
- The Authors certify that all authors have approved the publication of this manuscript and that there is no conflict of interest.
- The Authors confirm that the manuscript is their original work, has not received prior publication, is not under consideration for publication elsewhere, and has not been previously published.
- The Authors confirm that all authors listed on the title page have contributed significantly to the work, have read the manuscript, attest to the validity and legitimacy of the data and its interpretation, and agree to its submission.
- The Authors confirm that the manuscript is not copied from or plagiarized from any other published work.
- The Authors declare that the manuscript will not be submitted for publication in any other journal or magazine until a decision is made by the journal editors.
- If the manuscript is finally accepted for publication, the Authors confirm that they will either proceed with publication immediately or withdraw the manuscript in accordance with the journal’s withdrawal policies.
- The Authors agree that, upon publication of the manuscript in this journal, they transfer copyright or assign exclusive rights to the publisher, including commercial rights


