Predicting Student Academic Performance using Educational Data Mining and Feature Selection: Evidence from Higher Education

Authors:
Roshan Renji, R. Mahalakshmi

Addresses:
Department of Advanced Computing and Analytics, School of Computing Science, Vels Institute of Science, Technology and Advanced Studies, Chennai, Tamil Nadu, India.

Abstract:

Educational Data Mining (EDM) involves the extraction of valuable information and insights from educational data. EDM discerns patterns and trends in educational data, which may enhance academic curricula, pedagogical approaches, assessment techniques, and student performance. Recently, there has been renewed interest among EDM researchers in predicting students' academic performance based on factors such as entry-level attributes, family background, and educational accomplishments in the course. This study presents the findings of the EDM research conducted in the higher education institutions in Kerala, India. The data were collected from 1050 students studying their final year of a three-year course in various arts and science colleges in Kerala, India. The study focuses on developing data mining models to predict student achievement using their personal, pre-university, family, and other performance attributes. The comparison of performance accuracy when all the feature attributes are included and also with feature selection methods with reduced attributes revealed that the application of feature selection marginally enhances the performance of all the classifiers, Naive Bayes, SMO, IBK, JRip, and J48. The analysis demonstrated that the Relief ranking filter with the attribute ranking meta characterized the highest accuracy (79%) for the SMO classifier (25%). Correlation-based Feature Selection Greedy Stepwise search method produced the best prediction results for SMO and J48 classifiers (78.87%). Overall, results showed that the SMO classifier outperformed other standalone classifiers, such as Naïve Bayes, IBK, JRip, and J48, in terms of weighted average metrics such as Precision (0.760), Recall (0.789), Accuracy (0.788), F-measure (0.773), and Kappa statistic (0.604).

Keywords: Educational Data Mining; Prediction Techniques; Student Attributes; Feature Selection; Data Mining; Machine Learning; Artificial Neural Network; Generative Adversarial Network; Deep Learning.

Received on: 25/01/2025, Revised on: 26/03/2025, Accepted on: 03/06/2025, Published on: 05/06/2026

DOI: 10.69888/FTSTL.2026.000745

FMDB Transactions on Sustainable Techno Learning, 2026 Vol. 4 No. 2, Pages: 57-67

  • Views : 43
  • Downloads : 8
Download PDF