How accurate is a rebuilt control arm?
We built a reproducible benchmark to measure it, then published the result. Every number on this page is performance on that semi-synthetic benchmark, not a claim about a real trial.
Psych-ECA: A Reproducible Semi-Synthetic Benchmark for Synthetic Control Arms in Longitudinal Psychiatry
Aakash Bhagat, Shashank Choudhary, Sapien Labs
Submitted 7 July 2026
Three numbers, on a benchmark.
What the paper measured, stated plainly. These are results on a semi-synthetic benchmark where the true answer is known, not measurements from a real trial.
Counterfactual accuracy on the depression benchmark. Best or tied-best of the eight methods tested. RMSE is not comparable across instruments.
Semi-synthetic benchmarkEmpirical coverage of the nominal 90% prediction bands across the three conditions. Gradient boosting reaches 87 to 88%, uncalibrated OU 62 to 75%.
Semi-synthetic benchmarkReduction in control-arm bias from the inverse-intensity correction, when sicker patients are measured more often than others.
Semi-synthetic benchmarkThe control arm the trial never observes.
The paper's figures, redrawn as live charts. A benchmark can show the true counterfactual control outcome because it generates it. A real trial cannot, which is the whole problem this method addresses.
The control arm the trial never gets to observe
Recreated from Figure 1Bands that cover what they claim to cover
Recreated from Table 3Half the bias, once you correct for who gets measured
Recreated from Figure 3View benchmark resultsTables 2 to 4Hide tables
Every value below is a result on the semi-synthetic Psych-ECA benchmark, averaged over 8 seeds. Known counterfactual outcomes exist here because the benchmark generates them. They do not exist in a real trial, so this is not clinical validation.
Accuracy. RMSE and absolute treatment-effect bias in scale points, lower is better. Best or tied-best per column is marked. RMSE is not comparable across instruments because the scales differ.
| PHQ-9 | HAM-A | PANSS | ||||
|---|---|---|---|---|---|---|
| Method | RMSE | ATE bias | RMSE | ATE bias | RMSE | ATE bias |
| LOCF (carry forward) | 5.25 | 4.65 | 7.65 | 6.19 | 14.17 | 8.01 |
| Pooled RWD mean | 2.94 | 0.81 | 5.18 | 1.43 | 14.18 | 2.67 |
| kNN matching | 2.34 | 0.53 | 4.34 | 0.80 | 11.55 | 1.26 |
| Linear mixed model | 2.82 | 1.41 | 4.84 | 1.82 | 11.92 | 2.23 |
| Gradient boosting | 2.33 | 0.55 | 4.33 | 0.82 | 11.48 | 1.43 |
| OU bridge, naive | 2.47 | 1.05 | 4.51 | 1.67 | 11.49 | 2.48 |
| OU bridge, IIW | 2.31 | 0.61 | 4.29 | 0.95 | 11.31 | 1.44 |
| Scribe | 2.31 | 0.61 | 4.29 | 0.95 | 11.31 | 1.44 |
- RMSE values are not comparable across instruments. PHQ-9 runs 0 to 27, HAM-A 0 to 56, and PANSS 30 to 210, so a larger number on PANSS is not a worse result.
- Scribe and the inverse-intensity-weighted OU bridge report identical point results. The conformal layer in Scribe changes the prediction bands, not the point estimate.
- kNN matching reports slightly lower absolute ATE bias than Scribe in all three conditions (0.53 against 0.61, 0.80 against 0.95, 1.26 against 1.44). Point accuracy alone does not separate the methods.
Calibration. PICP is the share of true outcomes that fell inside the reported band; the target is 0.90. Width is in scale points. Values at or above the target are marked.
| PHQ-9 | HAM-A | PANSS | ||||
|---|---|---|---|---|---|---|
| Method | PICP | width | PICP | width | PICP | width |
| kNN matching | 0.863 | 7.2 | 0.857 | 13.2 | 0.850 | 35.1 |
| Linear mixed model | 0.940 | 10.4 | 0.934 | 17.6 | 0.955 | 47.3 |
| Gradient boosting | 0.877 | 7.3 | 0.870 | 13.2 | 0.873 | 35.7 |
| OU bridge, naive | 0.618 | 4.4 | 0.673 | 8.9 | 0.749 | 26.5 |
| Scribe | 0.955 | 9.2 | 0.935 | 15.8 | 0.927 | 40.6 |
- The linear mixed model also reaches nominal coverage, but with wider intervals and less accurate point estimates than Scribe.
- Scribe's bands are wider than gradient boosting and naive OU. That extra width is what buys coverage at or above the target, which the paper describes as slightly conservative as befits a regulatory setting.
- Widths stay separated by instrument because PHQ-9, HAM-A, and PANSS use different numerical ranges.
Trial decisions. Type-I error under the null (target 0.05) and power under the alternative, across none, moderate, and strong informative visit sampling. Each cell averages 3 conditions x 8 seeds = 24 reps per cell, so treat the rates as benchmark estimates. Type-I values at or below the target are marked.
| Type-I error (null) | Power (alt.) | |||||
|---|---|---|---|---|---|---|
| Method | none | moder. | strong | none | moder. | strong |
| LOCF (carry forward) | 0.958 | 0.958 | 0.958 | 1.000 | 1.000 | 1.000 |
| kNN matching | 0.542 | 0.042 | 0.000 | 0.958 | 0.958 | 0.958 |
| Linear mixed model | 0.500 | 0.458 | 0.542 | 0.875 | 0.958 | 0.958 |
| OU bridge, naive | 0.458 | 0.458 | 0.583 | 0.958 | 0.958 | 0.958 |
| Scribe | 0.500 | 0.167 | 0.000 | 0.958 | 0.958 | 0.917 |
- Scribe's false-positive rate falls as informative sampling strengthens, from 0.500 to 0.167 to 0.000. The moderate value of 0.167 is still above the nominal 0.05 target, so this is not nominal behaviour at every setting.
- kNN matching also reaches low Type-I error under moderate and strong informativeness (0.042 and 0.000). The paper's uniqueness claim is scoped to trajectory methods, not to every estimator.
- Each cell averages 24 replications, so one replication moves a rate by about 0.042. Treat these as benchmark estimates, not precise population rates.
- The paper does not disclose the numerical informativeness settings, and it does not disclose the treatment effect used under the alternative.
How the benchmark works.
Psych-ECA simulates patients whose symptom severity follows a mean-reverting process, standing in for spontaneous remission and placebo response. Each simulated patient carries both a treated trajectory and the control trajectory they would have followed without treatment, driven by the same underlying noise. That gives an exact per-patient counterfactual: the control outcome a real trial never gets to observe. The target is each treated patient's counterfactual control score at week 8.
Real-world donor records are sampled at informative times, so sicker patients are measured more often. That measurement pattern biases a naive control arm. An inverse-intensity correction reweights for it, and a weighted conformal layer turns each prediction into a calibrated band with measured coverage. Trial patients are measured at fixed protocol weeks; donor patients are measured at their informative clinical times.
Results are averaged over eight seeds, with 1,500 donor records and 300 trial patients, against eight estimators. The generator embeds a known dynamics and visit-intensity model, which favors correctly specified methods, and it does not capture every real-world complexity such as comorbidity, differential care, or instrument drift across sites. External validity has to be argued on real external control arms with sensitivity analysis. The generator, all eight estimators, the metrics, and the seeds are released with the paper.
The written record.
One preprint today. Where more work is still in preparation, we say so rather than imply more.
Selected for India Ascends 2026.
Sapien Labs was named as one of four startups in Lightspeed India’s India Ascends 2026 cohort. Coverage below, quoted as reported.
Named Sapien Labs as one of four startups in Lightspeed India’s India Ascends 2026 cohort.
Lightspeed selects 4 deeptech startups for its India Ascends 2026 cohortReported Sapien Labs among four deeptech ventures joining Lightspeed’s India Ascends platform.
Four Emerging Deeptech Startups Join Lightspeed’s India Ascends PlatformThe program provides a year-long strategic roadmap, access to Lightspeed’s network, and over $500,000 in cloud, AI, and software credits.
According to Lightspeed’s program announcementWhat stood out across this young cohort wasn’t just the technical depth, but the clarity and the intent to build enduring companies solving hard problems. These founders represent a new generation of builders who are thinking globally from day zero, and we are all in on their journeys.