Benchmarks

Accuracy & methodology

We test the pipeline the API runs in two ways: on real published figures, against the statistics their authors printed, and on simulated trials, where the true patient data are known exactly. None of the test figures were used in development.

The benchmark results could not be loaded. Please refresh the page.

The production pipeline, measured

The API reads each figure with Claude Opus 5.5: it finds every survival panel and reads the curves, the numbers-at-risk table and every printed statistic, zooming into the figure wherever the full view is too small to be sure. The Guyot (2012) algorithm then reconstructs patient-level data. Where our pixel measurement of a panel agrees with the model's reading, the measured curve replaces it. Every reconstruction is then checked against the statistics printed on the figure.

95%of printed hazard ratios reproduced within 10% (141 checks, real figures)
96%of printed medians reproduced within 10% (128 checks)
96%of 269 independent checks against printed values passed
91%of simulated figures with the HR within 5% of the true HR
0.09 momedian error in median survival (simulated)
100%of at-risk table cells read exactly (simulated)
154survival panels reconstructed from 40 published trials, whole figures in
100%of simulated figures with the true HR inside the reconstructed 95% CI

Real published figures

We ran the pipeline on 87 Kaplan–Meier figures from 40 randomised trials published open access (CC BY) and harvested automatically from PubMed Central. Each whole figure went in unedited, as rendered from the article PDF (75 figures; 12 from PMC's web images where no PDF match was found). The model found 161 survival panels and the pipeline reconstructed patient-level data for 154. Each reconstruction was compared with every statistic the authors printed on the figure or in its inset table.

Printed statisticChecksWithin tolerance*Median difference90th percentile
Hazard ratio14195%2.1%6.6%
Median survival12896%0.6%4.2%
Number of events†8698%0.0%2.4%
Number of patients†86100%0.0%0.0%

*10% for hazard ratios, medians and events; exact for N. Hazard ratios: 82% within 5%, 95% within 10%. †Consistency checks, not independent validation: when a figure prints N and event counts, the reconstruction uses them as inputs (as the Guyot method recommends), so agreement confirms the reading is internally consistent. Hazard ratios and medians are outputs of the reconstruction and are the real test. 123 of 135 panels passed every check; the rest are flagged in the API response so an analyst knows exactly which panels to review. Printed values are rounded, and printed hazard ratios are often stratified or adjusted, while the reconstruction gives an unstratified Cox estimate, so a few percent of difference is expected even from a perfect reading. Every validated panel is published, with its checks, in the trial library.

What the remaining differences are. The figures with a flagged check were read again, independently, by a larger model (Claude Fable 5.1). It reproduced 14 of the 14 flagged differences to within 5%. When two models agree on the reconstruction, the difference comes from the figure itself: printed hazard ratios that are adjusted or stratified, medians of arms with only 3 to 13 patients, where one event moves the median, and event counts that do not match the plotted curve. The API labels a flag that a second reading reproduced, so the analyst knows where to look.

Simulated figures with known truth

The same pipeline on 80 held-out simulated figures (seeds 11000–11079), where the true patient data are known. To isolate each part, the same model readings were also scored without pixel refinement ("model only"), and the measurement engine was run alone on the same figures.

MetricProductionModel onlyMeasurement only
Hazard ratio within 5% of truth91.2%93.8%85.5%
Hazard ratio within 10% of truth96.2%96.2%90.8%
Median |log HR| error0.0090.0070.012
Curve error (median of per-arm mean |ΔS|)0.00270.00270.0028
Error in median survival (median, months)0.090.080.10
Events per arm (median relative error)0.9%0.9%1.0%
At-risk table cells read exactly100.0%100.0%95.6%
Arm names read exactly95.9%95.9%88.1%
True HR inside reconstructed 95% CI100%100%100%

How these readings were produced. The model readings for this page were made by Claude Opus 5.5 running in Claude Code, with the production instructions and output schema, the same image the API would send, and the same zoom tool (identical crops, enlargement and ruler). Readers worked blind: they saw only the figure and its caption, never our results or the printed-value checks. The pipeline downstream of the reading is exactly the code the API runs. Through the API, the reading runs at a fixed effort level with the same tool and a budget of 14 zooms.

Measurement engine: 400 simulated figures

Pixel measurement and on-device OCR, without the vision model. It refines the model's curves in production and takes over when the model is unavailable. Results on the full held-out set of simulated figures:

Generated 25 September 2026 · engine v0.1.0 · 400 figures, seeds 7000–7399

All figures in the held-out set, including degraded, black-and-white and table-less figures. Hazard-ratio metrics use every figure where both of the first two arms were reconstructed (n = 357).

82%of hazard ratios within 5% of the true HR
90%of hazard ratios within 10% of the true HR
97%of figures where the true HR lies inside the reconstructed 95% CI
1.1%median hazard-ratio error (|log HR ratio|, as %)
0.08 momedian error in median survival
1.1%median error in the number of events per arm
0.003median mean |ΔS| between reconstructed and true curves
0.11%median RMST error, as a fraction of τ
92%of at-risk table cells read exactly
2.8%of at-risk cells misread (the rest are reported missing)
90%of arm names read exactly
99.5%of figures processed without an error

Distributions

MetricnMedianMean90th pctMax
|ΔlogHR| (per figure)3570.01130.08390.09194.2756
Median survival error (months, per arm)6380.0830.6140.65941.056
Median survival error (relative, per arm)6380.4%2.9%3.4%383.0%
Events error (relative, per arm)7861.1%9.3%26.1%328.2%
Curve mean |ΔS| (per arm)8240.00260.01450.00540.8268
Curve max |ΔS| (per arm)8240.0270.0560.0821.000
RMST error / τ (per arm)7860.11%0.87%0.41%70.28%
Processing time (s per figure)39811.51911.48415.89122.708

n counts figures for HR metrics and arms for the per-arm metrics. Means are pulled up by a small number of failures. That is why the page leads with medians and with the share of results inside fixed tolerances.

Breakdowns

The same metrics, split by what makes a figure hard. "Processed" is the share of figures that returned a result. HR columns only count figures where a hazard ratio could be reconstructed.

Numbers-at-risk table present or absent

Without a table, IPD is reconstructed from the true N per arm (N-only Guyot).

SubgroupFiguresProcessedHR nMedian |ΔlogHR|Mean |ΔlogHR|HR within 10%Median events errorAt-risk cells exact
no table (N supplied)5398%480.06580.098763%24.3%—
with at-risk table347100%3090.00950.081694%0.8%92%

Colour vs black & white

Black & white includes curves separated only by grey level or line style.

SubgroupFiguresProcessedHR nMedian |ΔlogHR|Mean |ΔlogHR|HR within 10%Median events errorAt-risk cells exact
black&white3197%230.04800.692952%5.3%73%
colour369100%3340.01100.041993%1.0%94%

Clean vs degraded images

Degraded = downscaled to 55–100% and/or JPEG quality 55–92.

SubgroupFiguresProcessedHR nMedian |ΔlogHR|Mean |ΔlogHR|HR within 10%Median events errorAt-risk cells exact
clean266100%2440.00930.023694%0.9%96%
degraded (downscaled/JPEG)13499%1130.01670.214180%1.7%85%

Shaded confidence bands

Figures with shaded 95% CI bands around each curve.

SubgroupFiguresProcessedHR nMedian |ΔlogHR|Mean |ΔlogHR|HR within 10%Median events errorAt-risk cells exact
CI bands59100%480.01760.255669%1.8%81%
no CI bands34199%3090.01040.057293%1.0%95%

Number of arms

SubgroupFiguresProcessedHR nMedian |ΔlogHR|Mean |ΔlogHR|HR within 10%Median events errorAt-risk cells exact
2 arms35499%3120.01150.072590%1.1%93%
3 arms46100%450.00960.162684%1.2%92%

Methodology

Simulated trials

Rendering

Held-out protocol

The engine was developed on separate seed ranges. The benchmark reported here uses seeds 7000–7399, generated after development and never used for tuning or threshold selection. Each figure goes through the same code path as a standard API request, with no hints. For figures without an at-risk table, reconstruction uses the true N per arm, as a user would supply it from the paper.

Metric definitions

HR within 5% / 10%
|log(HRreconstructed / HRtrue)| below log(1.05) or log(1.10), comparing arm 2 vs arm 1 on the reconstructed IPD with the same comparison on the true IPD.
|ΔlogHR|
Absolute difference in log hazard ratio. For small values it is roughly the relative HR error (0.01 ≈ 1%).
True HR in 95% CI
Whether the true HR lies inside the Wald 95% confidence interval computed from the reconstructed IPD.
Curve mean |ΔS|
Per arm: mean absolute difference between the reconstructed and true survival curves over 600 time points from 0 to the end of the plotted curve (an integrated absolute error normalised by time).
Median survival error
Per arm: absolute difference in months between the KM median of the reconstructed and true data, where both are reached within the plotted range.
Events error
Per arm: |eventsreconstructed − eventstrue| / eventstrue, counting true events within the plotted time range.
RMST error
Per arm: |RMSTreconstructed − RMSTtrue| / τ, where τ is the end of the plotted follow-up.
At-risk cells exact / misread
Share of numbers-at-risk cells, within the plotted range, returned with exactly the true value or with a wrong value. The remainder were not returned.
Arm names
Share of arms whose returned name matches the true name, ignoring case, spaces and punctuation.
Processed
Share of figures for which the engine returned a result instead of an error.

Known limits

Where results are least reliable, and what the API tells you when it happens.

The production pipeline

When a check against a printed value fails, the figure is read again by a larger model. If the second reading reproduces the difference, the check is labelled accordingly, which points to the figure rather than to a reading error.

The measurement engine (fallback)

Without the vision model, results come from pixel measurement and on-device OCR. It handles single panels only, and is weakest on:

Which model reads curves best?

We re-test whenever a new model ships. Each model below read the whole figure on its own, curves included, with no pixel measurement, and the same Guyot reconstruction ran on its reading. Scores are against the exact truth on 24 held-out simulated figures (seeds 11000–11023; the xhigh row uses the first 12). The first four rows are single API calls without the zoom tool.

Model (reading alone)Curve errorMedian-survival error (months)HR within 5%Median |log HR| errorCost / figure
Claude Opus 50.01550.7286%0.033~$0.06
Claude Fable 5.10.00750.4295%0.014~$0.14
Claude Opus 5.5 (high effort)0.00570.2995%0.011~$0.05
Claude Opus 5.5 (xhigh effort)0.00600.21100%0.008~$0.10
Claude Opus 5.5 with the zoom tool (production reader)0.00260.0792%0.006~$0.10–0.30

Curve error is the median of each arm's mean absolute difference in survival probability. With the zoom tool the model crops and enlarges the full-resolution figure wherever it needs to (the at-risk table, steps, crossings, the 50% line). That roughly halved the curve error and cut the median-survival error by three quarters, which matters more than effort level or a larger model. The zoom row was produced through Claude Code with the production instructions; API cost depends on how many zooms a figure needs.

How the pipeline got here

Earlier in 2026 the balance was the other way round. On the same 80 simulated figures, Claude Opus 5 reading curves on its own put 72.5% of hazard ratios within 5% of the truth, against 85.5% for pixel measurement, and its curve error was five times larger. So the API used a hybrid: the model read the text, and the curves were measured in pixels. Opus 5.5 with a zoom tool now reads curves as precisely as the measurement engine, and reads text far better, so the model's reading comes first. Pixel measurement checks each curve, replaces it where the two agree closely, and takes over when the model is unavailable.

How this compares

Published evaluations use different figure sets, mixes of difficulty and metric definitions, so this comparison is indicative only.

SourceSettingReported resultTrialCurve
Saluja et al. 2019Manual digitisation + Guyot, 118 curves from 55 RCTsMean HR error 0.0094Median |log HR| error 0.009 (simulated); median difference from printed HRs 2.1% (real figures)
Kim et al. 2025Reconstruction of 58 published HRs84% of HRs within 5%82% of 141 printed HRs within 5% (real figures); 91% of true HRs (simulated)
KM-GPT (2025)Automated pipeline, 540 synthetic plotsMedian absolute S(t) error 0.005Median per-arm mean |ΔS| 0.0027 (a related but not identical statistic)
ISPOR 2025 posterGPT-4o alone vs a dedicated computer-vision pipelineMedian survival error ±1.9 vs ±0.6 monthsMedian error in median survival 0.09 months (simulated); printed medians reproduced to a median 0.6% (real figures)

How to cite

TrialCurve (2026). Automated reconstruction of individual patient data from published Kaplan–Meier figures, engine v0.1.0. https://trialcurve.com/accuracy. Benchmark data: benchmark.json (CC BY 4.0).

Reproducing the benchmark

The figure generator and scoring harness are part of the TrialCurve codebase. Enterprise customers can get them on request for validation audits, together with the exact seed list, so that every number on this page can be regenerated.