Evidence

How accurate is a rebuilt control arm?

We built a reproducible benchmark to measure it, then published the result. Every number on this page is performance on that semi-synthetic benchmark, not a claim about a real trial.

arXiv:2607.27224v1Preprint, version 1Statistics (Applications), Machine Learning

Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry

Aakash Bhagat, Shashank Choudhary, Sapien Labs

Submitted 7 July 2026

01The result

Three numbers, on a benchmark.

What the paper measured, stated plainly. These are results on a semi-synthetic benchmark where the true answer is known, not measurements from a real trial.

2.3PHQ-9 RMSE, points

Counterfactual accuracy on the depression benchmark. Best or tied-best of the eight methods tested. RMSE is not comparable across instruments.

Semi-synthetic benchmark
93to96%Band coverage

Empirical coverage of the nominal 90% prediction bands across the three conditions. Gradient boosting reaches 87 to 88%, uncalibrated OU 62 to 75%.

Semi-synthetic benchmark
40to50%Bias removed

Reduction in control-arm bias from the inverse-intensity correction, when sicker patients are measured more often than others.

Semi-synthetic benchmark
02The figures, rebuilt

The control arm the trial never observes.

The paper's figures, redrawn as live charts. A benchmark can show the true counterfactual control outcome because it generates it. A real trial cannot, which is the whole problem this method addresses.

The control arm the trial never gets to observe

Recreated from Figure 1
Cohort mean, PHQ-9
8101214024681012
Weeks
Treated arm, observedTrue counterfactual controlScribe at week 8Naive bridge at week 8
Per patient at week 8, PHQ-9
0481216111223344
Treated patients, sorted
Scribe estimate90% calibrated bandTrue counterfactual
0.61Scribe absolute ATE bias at week 8, PHQ-9 points
1.05Naive OU bridge, same benchmark, same week
15.0Enrollment baseline, PHQ-9, before either arm moves
Semi-synthetic Psych-ECA benchmark, depression condition, PHQ-9 scale 0 to 27. The untreated control curve is the benchmark generator's own mean path from the paper's condition parameters, and both week-8 markers sit above it by exactly the absolute treatment-effect bias the paper reports for that method (1.05 naive, 0.61 Scribe). On the right, the band width of 9.2PHQ-9 points is the paper's measured mean width and 42 of 44 points fall inside it, matching the 95.5% coverage measured on this condition. Individual patient positions are an illustrative sample. Counterfactual truth exists here because the benchmark generates it. It does not exist in a real trial.

Bands that cover what they claim to cover

Recreated from Table 3
kNN matching, PHQ-9: 86.3% coveragekNN matching, HAM-A: 85.7% coveragekNN matching, PANSS: 85.0% coverageLinear mixed model, PHQ-9: 94.0% coverageLinear mixed model, HAM-A: 93.4% coverageLinear mixed model, PANSS: 95.5% coverageGradient boosting, PHQ-9: 87.7% coverageGradient boosting, HAM-A: 87.0% coverageGradient boosting, PANSS: 87.3% coverageOU bridge, naive, PHQ-9: 61.8% coverageOU bridge, naive, HAM-A: 67.3% coverageOU bridge, naive, PANSS: 74.9% coverageScribe, PHQ-9: 95.5% coverageScribe, HAM-A: 93.5% coverageScribe, PANSS: 92.7% coveragekNN matchingLinear mixed modelGradient boostingOU bridge, naiveScribeNominal 90%60708090100
Empirical coverage of the 90% band, percent
ScribeOther methodsNominal 90%Three dots per method, one per condition
Empirical coverage of the nominal 90% prediction band on the semi-synthetic Psych-ECA benchmark, one dot per condition. A method is calibrated when its coverage sits at or above the 90% line. Scribe reaches 95.5%, 93.5%, and 92.7% across the three conditions, described in the paper as slightly conservative as befits a regulatory setting. The linear mixed model also reaches nominal coverage, but with wider intervals and less accurate point estimates. Naive OU under-covers, at 61.8% to 74.9%. Coverage on a benchmark is not a guarantee for future datasets.

Half the bias, once you correct for who gets measured

Recreated from Figure 3
PHQ-9 · Depression
1.05
0.61
NaiveCorrected
42%bias removed
HAM-A · Anxiety
1.67
0.95
NaiveCorrected
43%bias removed
PANSS · Psychosis
2.48
1.44
NaiveCorrected
42%bias removed
Naive OU bridgeInverse-intensity corrected (Scribe)Scale points, per instrument, lower is better
Absolute external-control treatment-effect bias on the semi-synthetic Psych-ECA benchmark, in each instrument's own scale points. In the benchmark, sicker simulated patients are measured more often, which biases a naive control arm. Inverse-intensity weighting corrects for that measurement pattern and roughly halves the bias in every condition. Each panel is scaled to its own baseline; the scale points are not comparable across instruments. This is a benchmark result, not a measurement from a real trial.
View benchmark resultsTables 2 to 4Hide tables

Every value below is a result on the semi-synthetic Psych-ECA benchmark, averaged over 8 seeds. Known counterfactual outcomes exist here because the benchmark generates them. They do not exist in a real trial, so this is not clinical validation.

Seeds 8Donors 1,500Trial patients 300Visits weeks 0, 1, 2, 4, 6, 8, 10, 12Landmark Week 8Nominal band 90%Test 5% ECA z-test

Accuracy. RMSE and absolute treatment-effect bias in scale points, lower is better. Best or tied-best per column is marked. RMSE is not comparable across instruments because the scales differ.

PHQ-9HAM-APANSS
MethodRMSEATE biasRMSEATE biasRMSEATE bias
LOCF (carry forward)5.254.657.656.1914.178.01
Pooled RWD mean2.940.815.181.4314.182.67
kNN matching2.340.534.340.8011.551.26
Linear mixed model2.821.414.841.8211.922.23
Gradient boosting2.330.554.330.8211.481.43
OU bridge, naive2.471.054.511.6711.492.48
OU bridge, IIW2.310.614.290.9511.311.44
Scribe2.310.614.290.9511.311.44
  • RMSE values are not comparable across instruments. PHQ-9 runs 0 to 27, HAM-A 0 to 56, and PANSS 30 to 210, so a larger number on PANSS is not a worse result.
  • Scribe and the inverse-intensity-weighted OU bridge report identical point results. The conformal layer in Scribe changes the prediction bands, not the point estimate.
  • kNN matching reports slightly lower absolute ATE bias than Scribe in all three conditions (0.53 against 0.61, 0.80 against 0.95, 1.26 against 1.44). Point accuracy alone does not separate the methods.

Calibration. PICP is the share of true outcomes that fell inside the reported band; the target is 0.90. Width is in scale points. Values at or above the target are marked.

PHQ-9HAM-APANSS
MethodPICPwidthPICPwidthPICPwidth
kNN matching0.8637.20.85713.20.85035.1
Linear mixed model0.94010.40.93417.60.95547.3
Gradient boosting0.8777.30.87013.20.87335.7
OU bridge, naive0.6184.40.6738.90.74926.5
Scribe0.9559.20.93515.80.92740.6
  • The linear mixed model also reaches nominal coverage, but with wider intervals and less accurate point estimates than Scribe.
  • Scribe's bands are wider than gradient boosting and naive OU. That extra width is what buys coverage at or above the target, which the paper describes as slightly conservative as befits a regulatory setting.
  • Widths stay separated by instrument because PHQ-9, HAM-A, and PANSS use different numerical ranges.

Trial decisions. Type-I error under the null (target 0.05) and power under the alternative, across none, moderate, and strong informative visit sampling. Each cell averages 3 conditions x 8 seeds = 24 reps per cell, so treat the rates as benchmark estimates. Type-I values at or below the target are marked.

Type-I error (null)Power (alt.)
Methodnonemoder.strongnonemoder.strong
LOCF (carry forward)0.9580.9580.9581.0001.0001.000
kNN matching0.5420.0420.0000.9580.9580.958
Linear mixed model0.5000.4580.5420.8750.9580.958
OU bridge, naive0.4580.4580.5830.9580.9580.958
Scribe0.5000.1670.0000.9580.9580.917
  • Scribe's false-positive rate falls as informative sampling strengthens, from 0.500 to 0.167 to 0.000. The moderate value of 0.167 is still above the nominal 0.05 target, so this is not nominal behaviour at every setting.
  • kNN matching also reaches low Type-I error under moderate and strong informativeness (0.042 and 0.000). The paper's uniqueness claim is scoped to trajectory methods, not to every estimator.
  • Each cell averages 24 replications, so one replication moves a rate by about 0.042. Treat these as benchmark estimates, not precise population rates.
  • The paper does not disclose the numerical informativeness settings, and it does not disclose the treatment effect used under the alternative.
03Method

How the benchmark works.

Psych-ECA simulates patients whose symptom severity follows a mean-reverting process, standing in for spontaneous remission and placebo response. Each simulated patient carries both a treated trajectory and the control trajectory they would have followed without treatment, driven by the same underlying noise. That gives an exact per-patient counterfactual: the control outcome a real trial never gets to observe. The target is each treated patient's counterfactual control score at week 8.

Depression PHQ-9 0 to 27Anxiety HAM-A 0 to 56Psychosis PANSS 30 to 210

Real-world donor records are sampled at informative times, so sicker patients are measured more often. That measurement pattern biases a naive control arm. An inverse-intensity correction reweights for it, and a weighted conformal layer turns each prediction into a calibrated band with measured coverage. Trial patients are measured at fixed protocol weeks; donor patients are measured at their informative clinical times.

Results are averaged over eight seeds, with 1,500 donor records and 300 trial patients, against eight estimators. The generator embeds a known dynamics and visit-intensity model, which favors correctly specified methods, and it does not capture every real-world complexity such as comorbidity, differential care, or instrument drift across sites. External validity has to be argued on real external control arms with sensitivity analysis. The generator, all eight estimators, the metrics, and the seeds are released with the paper.

Press

Selected for India Ascends 2026.

Sapien Labs was named as one of four startups in Lightspeed India’s India Ascends 2026 cohort. Coverage below, quoted as reported.

What stood out across this young cohort wasn’t just the technical depth, but the clarity and the intent to build enduring companies solving hard problems. These founders represent a new generation of builders who are thinking globally from day zero, and we are all in on their journeys.
Hemant Mohapatra, Partner at Lightspeed