TensorX
返回文献探索

Paper · arXiv 2407.19340

Integrating Large Language Models into a Tri-Modal Architecture for Automated Depression Classification

Santosh V. Patapati

58 upvotesJuly 27, 2024arXiv 预印本
AI 摘要

A BiLSTM-based multi-modal model combines audio, facial, and text data using a GPT-4 model to classify depression, achieving state-of-the-art performance in binary classification tasks.

BiLSTMMel Frequency Cepstral CoefficientsFacial Action UnitsGPT-4two-shot learningmulti-modal architecture

Abstract

Major Depressive Disorder (MDD) is a pervasive mental health condition that affects 300 million people worldwide. This work presents a novel, BiLSTM-based tri-modal model-level fusion architecture for the binary classification of depression from clinical interview recordings. The proposed architecture incorporates Mel Frequency Cepstral Coefficients, Facial Action Units, and uses a two-shot learning based GPT-4 model to process text data. This is the first work to incorporate large language models into a multi-modal architecture for this task. It achieves impressive results on the DAIC-WOZ AVEC 2016 Challenge cross-validation split and Leave-One-Subject-Out cross-validation split, surpassing all baseline models and multiple state-of-the-art models. In Leave-One-Subject-Out testing, it achieves an accuracy of 91.01%, an F1-Score of 85.95%, a precision of 80%, and a recall of 92.86%.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号