SAD: Self-assessment of depression for Bangladeshi university students using machine learning and NLP

Depressive illness, influenced by social, psychological, and biological factors, is a significant public health concern that necessitates accurate and prompt diagnosis for effective treatment. This study explores the multifaceted nature of depression by investigating its correlation with various soc...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	Array (New York) 2025-03, Vol.25, p.100372, Article 100372
Hauptverfasser:	Azad, Md Shawmoon, Leeon, Shakirul Islam, Khan, Riasat, Mohammed, Nabeel, Momen, Sifat
Format:	Artikel
Sprache:	eng
Schlagworte:	BERT Chain of thought Depression diagnosis Explainable AI Large language models Machine learning Tree of thought
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Depressive illness, influenced by social, psychological, and biological factors, is a significant public health concern that necessitates accurate and prompt diagnosis for effective treatment. This study explores the multifaceted nature of depression by investigating its correlation with various social factors and employing machine learning, natural language processing, and explainable AI to analyze depression assessment scales. Data from a survey of 520 Bangladeshi university students, encompassing socio-personal and clinical questions, was utilized in this study. Eight machine learning algorithms with optimized hyperparameters were applied to evaluate eight depression assessment scales, identifying the most effective one. Additionally, ten machine learning models, including five BERT-based and two generative large language models, were tested using three prompting approaches and assessed across four categories of social factors: relationship dynamics, parental pressure, academic contentment, and exposure to violence. The results showed that support vector machines achieved a remarkable 99.14% accuracy with the PHQ9 scale. While considering the social factors, the stacking ensemble classifier demonstrated the best results. Among NLP approaches, BioBERT outperformed other BERT-based models with 90.34% accuracy when considering all social aspects. In prompting approaches, the Tree of Thought prompting on Claude Sonnet surpassed other prompting techniques with 75.00% accuracy. However, traditional machine learning models outshined NLP methods in tabular data analysis, with the stacking ensemble model achieving the highest accuracy of 97.88%. The interpretability of the top-performing classifier was ensured using LIME.
ISSN:	2590-0056 2590-0056
DOI:	10.1016/j.array.2024.100372