Real-world data in.
Trial-ready evidence out.
One configurable pipeline transforms real-world clinical data into external control cohorts, with provenance preserved at every step. It is therapeutic-area agnostic. The specifics per area live in the explorer.
Five steps, one governed path.
Each stage is auditable and hands the next a stronger artifact. The output is an evidence package aligned with external-control guidance.
- 01
Ingest and de-identify
Records are de-identified at source and ingested under strict protocols, inside the partner institution's governance boundary. Governance, access, and provenance are agreed with their team before any data moves.
De-identified at source. Nothing raw leaves the institution. - 02
Reconstruct patient journeys
Fragmented records across care settings are linked into coherent, longitudinal trajectories, still inside the boundary.
One patient, one timeline, assembled from scattered encounters. - 03
Derive and validate endpoints
Routine-care narratives, labs, and imaging are mapped to trial-standard structure and scored into validated endpoints with item-level traceability, checked against assessed ground truth with clinical partners.
Every scored item points back to the sentence, value, or read it came from. - 04
Assemble control arms
Protocol-aligned external cohorts are matched to your inclusion and exclusion criteria, and reused across studies rather than rebuilt for each one.
Matched to your protocol, not a fixed dataset. - 05
Hand off to your team
Only a de-identified evidence package crosses the boundary: regulatory-ready documentation aligned with external-control guidance, with every derived value traceable back for audit.
Your team leads the submission. We deliver the evidence they defend.
Every derived endpoint traces back to its source record
Every endpoint traces to a source record.
Select a derived endpoint and the exact sentence, measurement, or value it came from lights up in the source record. The same provenance holds across notes, imaging reads, and labs. All records shown are synthetic.
Every derived endpoint carries item-level provenance. Your biostatistics team can inspect any data point.
Follow-up. Pt reports low mood most days, little interest in usual activities, and reduced appetite. Sleep fragmented, waking around 4am, unable to return to sleep. Denies SI. Concentration poor at work. Affect constricted.
Item score derived from the sleep sentence, validated against assessed ground truth with clinical partners.
The whole engine, in one continuous move.
Scroll, and an unstructured record becomes a structured timeline, a validated endpoint, then one matched patient in a cohort that snaps into a control arm. This is the psychiatry case, our hardest and first: the outcome lives in prose. The same move runs across every therapeutic area.
- 01 / 04Note
A real-looking clinical narrative. The outcome lives in prose, not a structured field.
- 02 / 04Timeline
Fragmented records link into one longitudinal trajectory across care settings.
- 03 / 04Endpoint
Each scored item traces back to the exact sentence it was derived from, checked against assessed ground truth.
- 04 / 04Cohort
The validated patient joins a protocol-aligned cohort. One glyph marks the record this still followed.
Illustrative, built from a synthetic note. No real patient record is shown.
Every stage runs inside the boundary.
Raw records never move. The pipeline runs where the data already lives, and the only thing that crosses out is a de-identified, audit-traceable evidence package.
Governance, access, and provenance are defined with the partner institution before any data moves. Raw records stay inside; only the de-identified evidence package crosses out.