Artificial intelligence has entered psychological services through diagnostic classifiers, clinical-language models, passive sensing, conversational treatment, virtual-reality practice, and remote psychotherapy. Evidence for these uses is commonly discussed by technical family, although a clinical decision also depends on how directly a system reaches a patient, whether the study design permits causal interpretation, whether the endpoint concerns symptoms rather than an intermediate signal, and whether the population size can be audited. This study asks which of those properties governs the robustness of clinical evidence proximity across early applications in psychological care. The analytic cohort comprised nine investigations published between 2011 and 2020. Seven provided exact participant counts, representing 20,116 people; one of these also analysed 90,934 psychotherapy transcripts. Each record was coded for population scale, design strength, care proximity, and endpoint proximity. A weighted clinical evidence proximity score was examined with 100,000 Dirichlet weight perturbations, probabilistic treatment of two unreported participant totals, complete criterion omission, and an exact 126-allocation permutation test. Median scores ranged from 42.4 to 85.4. Automated cognitive behavioural therapy for insomnia and telephone psychotherapy occupied first place in 43.1% and 48.9% of perturbations, respectively, whereas conversational cognitive behavioural therapy entered the first three positions in 99.95%. Patient-facing interventions exceeded the remaining applications by 19.2 points on a care-excluded three-criterion score (exact p = 0.0238). Large population size alone did not secure a leading position: the 17,572-patient psychotherapy-language study had a median score of 64.9 because it measured treatment process rather than symptom change. The research question is therefore answered by a joint condition. Direct care delivery was associated with higher evidentiary proximity, but stable leadership required both a comparative design and an endpoint connected to clinical change. The findings support transparent, criterion-specific assessment when psychological artificial intelligence is considered for clinical use.
Artificial intelligence has entered psychological services through diagnostic classifiers, clinical-language models, passive sensing, conversational treatment, virtual-reality practice, and remote psychotherapy. Evidence for these uses is commonly discussed by technical family, although a clinical decision also depends on how directly a system reaches a patient, whether the study design permits causal interpretation, whether the endpoint concerns symptoms rather than an intermediate signal, and whether the population size can be audited. This study asks which of those properties governs the robustness of clinical evidence proximity across early applications in psychological care. The analytic cohort comprised nine investigations published between 2011 and 2020. Seven provided exact participant counts, representing 20,116 people; one of these also analysed 90,934 psychotherapy transcripts. Each record was coded for population scale, design strength, care proximity, and endpoint proximity. A weighted clinical evidence proximity score was examined with 100,000 Dirichlet weight perturbations, probabilistic treatment of two unreported participant totals, complete criterion omission, and an exact 126-allocation permutation test. Median scores ranged from 42.4 to 85.4. Automated cognitive behavioural therapy for insomnia and telephone psychotherapy occupied first place in 43.1% and 48.9% of perturbations, respectively, whereas conversational cognitive behavioural therapy entered the first three positions in 99.95%. Patient-facing interventions exceeded the remaining applications by 19.2 points on a care-excluded three-criterion score (exact p = 0.0238). Large population size alone did not secure a leading position: the 17,572-patient psychotherapy-language study had a median score of 64.9 because it measured treatment process rather than symptom change. The research question is therefore answered by a joint condition. Direct care delivery was associated with higher evidentiary proximity, but stable leadership required both a comparative design and an endpoint connected to clinical change. The findings support transparent, criterion-specific assessment when psychological artificial intelligence is considered for clinical use.