Optimizing Treatment Strategies in the Bipolar Disorder Spectrum With Classical AI Approaches: Systematic Review of Performance, Bias, and Clinical Applicability

Scritto il 21/07/2026
da Silvia De Francesco

JMIR Ment Health. 2026 Jul 21;13:e93307. doi: 10.2196/93307.

ABSTRACT

BACKGROUND: Bipolar disorder (BD) is a complex and heterogeneous psychiatric condition, characterized by fluctuating clinical courses that affect approximately 1%-2% of the global population in their lifetime. Despite pharmacological advances, treatment response varies significantly among patients, making the identification of individualized treatment strategies a major challenge. Artificial Intelligence (AI), through its classical approaches, has emerged as a powerful tool in precision psychiatry to identify subtle patterns in complex data and inform personalized clinical decisions.

OBJECTIVE: The present systematic review aimed to examine the current evidence on classical AI-supported treatment optimization in the BD spectrum.

METHODS: The review was conducted in accordance with the PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) 2020 guidelines. Four databases (PubMed, Web of Science, Scopus, and Embase) were searched for original studies published after 2015 on the application of classical AI in the treatment of BD in adult patients. Publication bias was evaluated by visual inspection of a funnel plot. The methodological quality, risk of bias, and clinical applicability of the predictive models were assessed using the Prediction Model Risk Of Bias Assessment Tool for prediction models using regression or AI methods (PROBAST+AI; PROBAST+AI Working Group) tool.

RESULTS: A total of 35 studies were included and classified into 5 outcome-based categories, including acute symptomatic response, long-term maintenance response, relapse and readmission risk, safety and dose optimization, and brain aging and phenotyping. Acute symptomatic response models performed modestly (pooled area under the curve [AUC] 0.68), while imaging improved accuracy (74%-77%). Long-term maintenance response models showed moderate-to-high performance (pooled AUC 0.80), with biomarker- and cellular-based models reaching 96%-99% accuracy. Relapse and readmission prediction achieved a pooled AUC of 0.71, with digital phenotyping and rule-based methods performing best (AUC 0.85-0.88). Safety and dose optimization models achieved 85%-97% accuracy. Brain aging and phenotyping studies highlighted accelerated brain aging in BD, partially mitigated by lithium, and revealed novel data-driven subgroups. However, 3 studies were considered at high risk of bias due to small sample sizes associated with disproportionately high-performance estimates. An additional study was identified as potentially biased because it lay markedly distant from the funnel plot's confidence line. Finally, the PROBAST+AI assessment revealed a high risk of bias in most studies, primarily due to data analysis limitations, small sample sizes, and lack of external validation.

CONCLUSIONS: The adoption of classical AI tools in BD serves as a driver for therapeutic optimization, although current AI tools in BD should still be considered exploratory rather than ready for clinical use. Effective implementation in real-world clinical scenarios requires more robust, transparent, and externally validated models to ensure reliability and generalizability.

PMID:42479879 | DOI:10.2196/93307