When Transportability Matters More Than Complexity: Multicentre Calibration of a High-Frequency rTMS Response Score in an In-Silico Population Study

Barbara H. Stanley1
1Department of Psychiatry, Vagelos College of Physicians and Surgeons, Columbia University, New York, USA
Published: 11/10/2024
Cite this article as: Barbara H. Stanley. When Transportability Matters More Than Complexity: Multicentre Calibration of a High-Frequency rTMS Response Score in an In-Silico Population Study. Psilogos, Issue 1. Page: 1-11.

Abstract

Prognostic models for repetitive transcranial magnetic stimulation (rTMS) are commonly judged by discrimination within the population from which they were estimated. Clinical deployment instead requires accurate probabilities after the diagnostic composition, symptom severity, treatment resistance, and anatomical characteristics of the treated population have changed. We examined whether a site-aware penalized recalibration procedure improves transport of a published high-frequency left dorsolateral prefrontal cortex response score and whether added complexity is justified relative to a centre-specific intercept correction. Eight prespecified in-silico populations contained 48,000 records with systematic differences in sex composition, bipolar-depression prevalence, age, Montgomery–Åsberg Depression Rating Scale score, treatment-resistance stage, dorsolateral prefrontal cortical thickness, and scalp-to-cortex distance. Outcomes were sampled from a nonlinear probability function anchored to published coefficients but altered by centre and case mix. Four approaches were compared: the unchanged transported score, centre-specific intercept correction, pooled logistic recalibration, and partially pooled transport calibration (PPTC). Four hundred records per centre were reserved for updating and 5,600 per centre for held-out evaluation. Probability accuracy, calibration, discrimination, and decision-curve utility were quantified. Paired differences were evaluated by 600 centre-stratified bootstrap samples. Among 44,800 held-out records, response prevalence was 45.24%. The unchanged score attained an area under the receiver-operating-characteristic curve of 0.790 and a Brier score of 0.1856, yet its largest centre-level absolute calibration error was 9.90 percentage points. Centre-specific intercept correction reduced that maximum error to 2.46 points, improved the Brier score to 0.1812, and increased net benefit at a 0.35 treatment threshold from 0.2498 to 0.2579. Corresponding paired differences were −0.00434 (95% bootstrap interval −0.00496 to −0.00377) for the Brier score and +0.00809 (0.00626 to 0.00997) for net benefit. PPTC reduced the maximum centre error to 5.61 points and improved net benefit to 0.2567, but it did not outperform intercept correction; its largest subgroup error was 2.92 points. Pooled recalibration offered negligible improvement. The research question was whether a site-aware multivariable update warrants its complexity when an rTMS response model crosses populations. Under the prespecified case-mix changes, it did not. A local intercept correction was the most reliable approach because it addressed prevalence drift without destabilizing individual risk contrasts. High discrimination alone concealed clinically important centre-level probability errors; local calibration should therefore precede threshold-based use of transported rTMS response scores.

Keywords: repetitive transcranial magnetic stimulation; treatment-resistant depression; probability calibration; transportability; decision curve; structural magnetic resonance imaging; in-silico study

Abstract

Prognostic models for repetitive transcranial magnetic stimulation (rTMS) are commonly judged by discrimination within the population from which they were estimated. Clinical deployment instead requires accurate probabilities after the diagnostic composition, symptom severity, treatment resistance, and anatomical characteristics of the treated population have changed. We examined whether a site-aware penalized recalibration procedure improves transport of a published high-frequency left dorsolateral prefrontal cortex response score and whether added complexity is justified relative to a centre-specific intercept correction. Eight prespecified in-silico populations contained 48,000 records with systematic differences in sex composition, bipolar-depression prevalence, age, Montgomery–Åsberg Depression Rating Scale score, treatment-resistance stage, dorsolateral prefrontal cortical thickness, and scalp-to-cortex distance. Outcomes were sampled from a nonlinear probability function anchored to published coefficients but altered by centre and case mix. Four approaches were compared: the unchanged transported score, centre-specific intercept correction, pooled logistic recalibration, and partially pooled transport calibration (PPTC). Four hundred records per centre were reserved for updating and 5,600 per centre for held-out evaluation. Probability accuracy, calibration, discrimination, and decision-curve utility were quantified. Paired differences were evaluated by 600 centre-stratified bootstrap samples. Among 44,800 held-out records, response prevalence was 45.24%. The unchanged score attained an area under the receiver-operating-characteristic curve of 0.790 and a Brier score of 0.1856, yet its largest centre-level absolute calibration error was 9.90 percentage points. Centre-specific intercept correction reduced that maximum error to 2.46 points, improved the Brier score to 0.1812, and increased net benefit at a 0.35 treatment threshold from 0.2498 to 0.2579. Corresponding paired differences were −0.00434 (95% bootstrap interval −0.00496 to −0.00377) for the Brier score and +0.00809 (0.00626 to 0.00997) for net benefit. PPTC reduced the maximum centre error to 5.61 points and improved net benefit to 0.2567, but it did not outperform intercept correction; its largest subgroup error was 2.92 points. Pooled recalibration offered negligible improvement. The research question was whether a site-aware multivariable update warrants its complexity when an rTMS response model crosses populations. Under the prespecified case-mix changes, it did not. A local intercept correction was the most reliable approach because it addressed prevalence drift without destabilizing individual risk contrasts. High discrimination alone concealed clinically important centre-level probability errors; local calibration should therefore precede threshold-based use of transported rTMS response scores.

Keywords: repetitive transcranial magnetic stimulation; treatment-resistant depression; probability calibration; transportability; decision curve; structural magnetic resonance imaging; in-silico study
Barbara H. Stanley
Department of Psychiatry, Vagelos College of Physicians and Surgeons, Columbia University, New York, USA

DOI

Cite this article as:

Barbara H. Stanley. When Transportability Matters More Than Complexity: Multicentre Calibration of a High-Frequency rTMS Response Score in an In-Silico Population Study. Psilogos, Issue 1. Page: 1-11.

Publication history

Copyright © 2026 Barbara H. Stanley. This is an open access article distributed under the Creative Commons Attribution License, which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.

Browse Advance Search