DeltaBench
Open standard · proposal stage

How close is the AI draft to the final spec?

DeltaBench measures how far AI-generated specifications, programs and outputs deviate from their approved versions, and classifies every change by what went wrong and how much it matters.

Four layers of measurement

Comparing text is not enough. DeltaBench checks whether the logic, the data and the tables agree, and what it cost a person to get there.

01

Specifications

SDTM and ADaM specs compared variable by variable.

  • Recall and precision
  • Attribute agreement
  • Derivation equivalence
02

Programs

SAS, R and Python compared by structure and by result.

  • Code retained
  • Same input, same output?
  • QC iterations
03

Outputs

Tables checked against their final values.

  • Cell-level match
  • Shell conformance
  • Population errors
04

Effort

What it took to reach the approved version.

  • Edit time
  • Review cycles
  • Versus manual build

Every change gets a label and a weight

Style edits and protocol amendments are not tool errors. The taxonomy separates real mistakes from normal study evolution, and weights them by consequence.

Hallucination

A variable, codelist or rule that does not exist.

high
Wrong derivation

Logic that produces different values.

high
Omission

Something required that was left out.

medium
Terminology error

Wrong controlled terminology or codelist value.

medium
Conformance issue

A change that introduces a new validation finding.

medium
Style only

Wording, order or formatting with no effect on data.

not counted
Upstream change

Caused by a protocol or SAP amendment, not the tool.

excluded

A public benchmark with known answers

Final versions can contain errors too. So tools are also scored on synthetic Synthebo studies, where the correct answer is known exactly.

  • Hidden test setsScored studies are regenerated each cycle, so they cannot be memorised.
  • Vendors run their own toolsWith their own models and keys; DeltaBench only scores the outputs.
  • Automated and human-in-the-loopRanked separately, next to a human programmer baseline.
EntryWeighted scoreSpec recallProgram accuracy
Human baseline0.91
97%99%
Tool A0.78
93%88%
Tool B0.71
90%84%
Tool C0.64
86%79%
Sample layout only. No tools have been scored yet; first results are planned for 2027.

Sponsors keep their data

The calculator runs inside your company. Specs, code and data never leave. Only anonymous totals are shared, if you choose to contribute to the industry report.

1. Capture

A small hook saves each AI draft the moment it is generated.

2. Score locally

DeltaBench compares draft and final on your systems and produces a scorecard.

3. Share totals (optional)

Anonymous aggregates feed a yearly industry report.

Roadmap

Definitions first, then tools, then results.

  1. Metric definitionsWhite paper and taxonomy for public comment
  2. CalculatorOpen-source scoring for specs, then programs
  3. Public benchmarkHidden Synthebo test sets and leaderboard
  4. Sponsor pilotsTwo or three organizations, run locally
  5. Industry reportFirst federated results across sponsors

Get involved

DeltaBench works only if the definitions are shared and neutral. Help shape them.

Programmers and statisticians

Review the taxonomy and help label reference examples.

Sponsors and CROs

Pilot the calculator on your own studies, with no data leaving your company.

Tool builders

Test against the public practice set and enter the first benchmark cycle.