Research28 Sep 2026That day's edition
ProgramDistill, a benchmark built from 26 web apps, scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8.
Built from 26 functioning reference applications, the benchmark scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8 percent on full-app reconstruction.
Check our sources · 1 source, 3 claimsYour reaction
One press adds one. We count a number for each notice and day, never who pressed it.
Check our sources
1 source, 3 claims. We opened the source and checked every sentence above against it.
1 ProgramDistill: benchmarking coding agents on feature discovery in real applications
Open the source-
The benchmark draws on 26 fully functional reference web applications, from which the researchers extracted 1,975 replay-verified behaviors and built 4,063 tasks.
26 applications ... 1,975 replay-verified behaviors ... 4,063 tasks without human intervention
-
On cumulative full-application reconstruction workflows, GPT-6 Astra achieved 49.2 percent success while Claude Opus 5 reached 28.8 percent.
GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows
-
The paper was submitted to arXiv on September 16, 2026.
Submitted on 16 Sep 2026
Nothing appears on this site that we have not opened and linked.