ResearchPublished All news from that day
Ongoing case: OpenAI's models Outside test · 14 storiesSdkb / Selena Deckelmann at Wikimania 2025 / CC BY-SA 4.0, croppedProgramDistill, a benchmark built from 26 web apps, scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8.
Built from 26 functioning reference applications, the benchmark scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8 percent on full-app reconstruction.
Check our sources · 3 facts from 1 sourceYour reaction
We count reactions per story and day, never who reacted. The counts help us choose what goes in the monthly issue. If you are signed in, your own page shows yours too.
Check our sources
Every sentence above is checked against this source.
1 ProgramDistill: benchmarking coding agents on feature discovery in real applications
Open the source-
The benchmark draws on 26 fully functional reference web applications, from which the researchers extracted 1,975 replay-verified behaviors and built 4,063 tasks.
26 applications ... 1,975 replay-verified behaviors ... 4,063 tasks without human intervention
-
On cumulative full-application reconstruction workflows, GPT-6 Astra achieved 49.2 percent success while Claude Opus 5 reached 28.8 percent.
GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows
-
The paper was submitted to arXiv on September 16, 2026.
Submitted on 16 Sep 2026
Topics
The morning email
On the mornings we publish: the three top stories and up to four short ones. Free.
We email you a link to confirm. An issue may include one sponsor, always labelled Sponsored · Advertisement. Our emails count opens and clicks, not who made them. Unsubscribe in one click. What we keep
