Auditability and Care Proximity in Artificial Intelligence for Psychological Services: A Study-Level Robustness Analysis

Jongsuk Ro1
1Department of Psychiatry, College of Medicine, Seoul National University, Seoul, South Korea
Published: 11/10/2024
Cite this article as: Jongsuk Ro. Auditability and Care Proximity in Artificial Intelligence for Psychological Services: A Study-Level Robustness Analysis. Psilogos, Issue 1. Page: 35-45.

Abstract

Artificial intelligence has entered psychological services through diagnostic classifiers, clinical-language models, passive sensing, conversational treatment, virtual-reality practice, and remote psychotherapy. Evidence for these uses is commonly discussed by technical family, although a clinical decision also depends on how directly a system reaches a patient, whether the study design permits causal interpretation, whether the endpoint concerns symptoms rather than an intermediate signal, and whether the population size can be audited. This study asks which of those properties governs the robustness of clinical evidence proximity across early applications in psychological care. The analytic cohort comprised nine investigations published between 2011 and 2020. Seven provided exact participant counts, representing 20,116 people; one of these also analysed 90,934 psychotherapy transcripts. Each record was coded for population scale, design strength, care proximity, and endpoint proximity. A weighted clinical evidence proximity score was examined with 100,000 Dirichlet weight perturbations, probabilistic treatment of two unreported participant totals, complete criterion omission, and an exact 126-allocation permutation test. Median scores ranged from 42.4 to 85.4. Automated cognitive behavioural therapy for insomnia and telephone psychotherapy occupied first place in 43.1% and 48.9% of perturbations, respectively, whereas conversational cognitive behavioural therapy entered the first three positions in 99.95%. Patient-facing interventions exceeded the remaining applications by 19.2 points on a care-excluded three-criterion score (exact p = 0.0238). Large population size alone did not secure a leading position: the 17,572-patient psychotherapy-language study had a median score of 64.9 because it measured treatment process rather than symptom change. The research question is therefore answered by a joint condition. Direct care delivery was associated with higher evidentiary proximity, but stable leadership required both a comparative design and an endpoint connected to clinical change. The findings support transparent, criterion-specific assessment when psychological artificial intelligence is considered for clinical use.

Keywords: clinical artificial intelligence; psychological intervention; diagnostic classification; digital mental health; evidence audit; sensitivity analysis

Abstract

Artificial intelligence has entered psychological services through diagnostic classifiers, clinical-language models, passive sensing, conversational treatment, virtual-reality practice, and remote psychotherapy. Evidence for these uses is commonly discussed by technical family, although a clinical decision also depends on how directly a system reaches a patient, whether the study design permits causal interpretation, whether the endpoint concerns symptoms rather than an intermediate signal, and whether the population size can be audited. This study asks which of those properties governs the robustness of clinical evidence proximity across early applications in psychological care. The analytic cohort comprised nine investigations published between 2011 and 2020. Seven provided exact participant counts, representing 20,116 people; one of these also analysed 90,934 psychotherapy transcripts. Each record was coded for population scale, design strength, care proximity, and endpoint proximity. A weighted clinical evidence proximity score was examined with 100,000 Dirichlet weight perturbations, probabilistic treatment of two unreported participant totals, complete criterion omission, and an exact 126-allocation permutation test. Median scores ranged from 42.4 to 85.4. Automated cognitive behavioural therapy for insomnia and telephone psychotherapy occupied first place in 43.1% and 48.9% of perturbations, respectively, whereas conversational cognitive behavioural therapy entered the first three positions in 99.95%. Patient-facing interventions exceeded the remaining applications by 19.2 points on a care-excluded three-criterion score (exact p = 0.0238). Large population size alone did not secure a leading position: the 17,572-patient psychotherapy-language study had a median score of 64.9 because it measured treatment process rather than symptom change. The research question is therefore answered by a joint condition. Direct care delivery was associated with higher evidentiary proximity, but stable leadership required both a comparative design and an endpoint connected to clinical change. The findings support transparent, criterion-specific assessment when psychological artificial intelligence is considered for clinical use.

Keywords: clinical artificial intelligence; psychological intervention; diagnostic classification; digital mental health; evidence audit; sensitivity analysis
Jongsuk Ro
Department of Psychiatry, College of Medicine, Seoul National University, Seoul, South Korea

DOI

Cite this article as:

Jongsuk Ro. Auditability and Care Proximity in Artificial Intelligence for Psychological Services: A Study-Level Robustness Analysis. Psilogos, Issue 1. Page: 35-45.

Publication history

Copyright © 2026 Jongsuk Ro. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

Browse Advance Search