Today ainotis Join

My notis

ResearchPublished All news from that day

Ongoing case: OpenAI's models Outside test · 14 storiesSdkb / Selena Deckelmann at Wikimania 2025 / CC BY-SA 4.0, cropped

ProgramDistill, a benchmark built from 26 web apps, scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8.

Built from 26 functioning reference applications, the benchmark scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8 percent on full-app reconstruction.

Share

Check our sources · 3 facts from 1 source

Share this story

Your reaction

We count reactions per story and day, never who reacted. The counts help us choose what goes in the monthly issue. If you are signed in, your own page shows yours too.

Check our sources

Every sentence above is checked against this source.

1 ProgramDistill: benchmarking coding agents on feature discovery in real applicationsarXiv · 16 Sep 2026 · 3 facts Open the source
  1. The benchmark draws on 26 fully functional reference web applications, from which the researchers extracted 1,975 replay-verified behaviors and built 4,063 tasks.

    26 applications ... 1,975 replay-verified behaviors ... 4,063 tasks without human intervention
  2. On cumulative full-application reconstruction workflows, GPT-6 Astra achieved 49.2 percent success while Claude Opus 5 reached 28.8 percent.

    GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows
  3. The paper was submitted to arXiv on September 16, 2026.

    Submitted on 16 Sep 2026

Topics

The morning email

On the mornings we publish: the three top stories and up to four short ones. Free.

We email you a link to confirm. An issue may include one sponsor, always labelled Sponsored · Advertisement. Our emails count opens and clicks, not who made them. Unsubscribe in one click. What we keep