A Comparative Analysis of Machine Learning models in the Classification Techniques in the Prediction of Autism in Children

No Thumbnail Available

Date

2025-12

Journal Title

Journal ISSN

Volume Title

Publisher

Lead City University, Ibadan

Abstract

Autism Spectrum Disorder (ASD) is a mild cognitive impairment known to affect a person’s language, communication, thought process and social behaviour. Some recent advances in machine learning made it possible to predict ASD using behavioural, demographic, and medical data, but the choice of optimal algorithms and feature selection techniques remains an open research question. This study used the ASD for children dataset obtained from the University of California, Irvine repository to investigate the performance of the selected machine learning models through a comparative analysis. Eight nonparametric models were compared in total, with four models SVM, KNN, GNB, and MLP being base learners, and the other four ensemble methods RF, XGBoost, Bootstrap Aggregating (Bagging), and Stacking. Hyperparameters of the models were autotuned, and GridSearchCV was employed to choose the model with the best combination of parameters. All base models performed brilliantly, with SVM being the best and most precise in this experiment with a result of 1, amongst all classifiers even before hyperparameter tuning. After hyperparameter tuning, SVM and MLP ranked highest with all metrics returning scores of 100%. KNN returned 96.2% accuracy and recall and GNB trailed at 95.26% for both metrics for the ensemble methods, XGB and Bagging with SVM performed best at 100% for all metrics. RF produced an accuracy and recall of 98.1% while Bagging with KNN and GNB had both metrics at 95.2%. The classification performance of Bagging with MLP and the stacking classifier were equivalent across all metrics though, producing accuracy and recall of 99.5%. Based on the findings, it is recommended that thorough feature selection should be conducted to eliminate features that could lead to overfitting, such as "Q_Chat_10" in this study cross-validation technique was used to assess the generalizability of models to unseen data. Keywords: Accuracy, Classifier, Defaults, Financial, Models, Predicting, Validation Word Count: 300

Description

Keywords

Accuracy, Classifier, Defaults, Financial, Models, Predicting, Validation

Citation

Kate Turabian