How close is the AI draft to the final spec?
DeltaBench measures how far AI-generated specifications, programs and outputs deviate from their approved versions, and classifies every change by what went wrong and how much it matters.
Four layers of measurement
Comparing text is not enough. DeltaBench checks whether the logic, the data and the tables agree, and what it cost a person to get there.
Specifications
SDTM and ADaM specs compared variable by variable.
- Recall and precision
- Attribute agreement
- Derivation equivalence
Programs
SAS, R and Python compared by structure and by result.
- Code retained
- Same input, same output?
- QC iterations
Outputs
Tables checked against their final values.
- Cell-level match
- Shell conformance
- Population errors
Effort
What it took to reach the approved version.
- Edit time
- Review cycles
- Versus manual build
Every change gets a label and a weight
Style edits and protocol amendments are not tool errors. The taxonomy separates real mistakes from normal study evolution, and weights them by consequence.
A variable, codelist or rule that does not exist.
highLogic that produces different values.
highSomething required that was left out.
mediumWrong controlled terminology or codelist value.
mediumA change that introduces a new validation finding.
mediumWording, order or formatting with no effect on data.
not countedCaused by a protocol or SAP amendment, not the tool.
excludedA public benchmark with known answers
Final versions can contain errors too. So tools are also scored on synthetic Synthebo studies, where the correct answer is known exactly.
- Hidden test setsScored studies are regenerated each cycle, so they cannot be memorised.
- Vendors run their own toolsWith their own models and keys; DeltaBench only scores the outputs.
- Automated and human-in-the-loopRanked separately, next to a human programmer baseline.
| Entry | Weighted score | Spec recall | Program accuracy | |
|---|---|---|---|---|
| Human baseline | 0.91 | 97% | 99% | |
| Tool A | 0.78 | 93% | 88% | |
| Tool B | 0.71 | 90% | 84% | |
| Tool C | 0.64 | 86% | 79% |
Sponsors keep their data
The calculator runs inside your company. Specs, code and data never leave. Only anonymous totals are shared, if you choose to contribute to the industry report.
1. Capture
A small hook saves each AI draft the moment it is generated.
2. Score locally
DeltaBench compares draft and final on your systems and produces a scorecard.
3. Share totals (optional)
Anonymous aggregates feed a yearly industry report.
Roadmap
Definitions first, then tools, then results.
- Metric definitionsWhite paper and taxonomy for public comment
- CalculatorOpen-source scoring for specs, then programs
- Public benchmarkHidden Synthebo test sets and leaderboard
- Sponsor pilotsTwo or three organizations, run locally
- Industry reportFirst federated results across sponsors
Get involved
DeltaBench works only if the definitions are shared and neutral. Help shape them.
Programmers and statisticians
Review the taxonomy and help label reference examples.
Sponsors and CROs
Pilot the calculator on your own studies, with no data leaving your company.
Tool builders
Test against the public practice set and enter the first benchmark cycle.