Today ainotis

Research28 Sep 2026That day's edition

ProgramDistill, a benchmark built from 26 web apps, scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8.

Built from 26 functioning reference applications, the benchmark scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8 percent on full-app reconstruction.

Check our sources · 1 source, 3 claims

Your reaction

One press adds one. We count a number for each notice and day, never who pressed it.

Check our sources

1 source, 3 claims. We opened the source and checked every sentence above against it.

1 ProgramDistill: benchmarking coding agents on feature discovery in real applicationsarXiv · 16 Sep 2026 · 3 claims Open the source
  1. The benchmark draws on 26 fully functional reference web applications, from which the researchers extracted 1,975 replay-verified behaviors and built 4,063 tasks.

    26 applications ... 1,975 replay-verified behaviors ... 4,063 tasks without human intervention
  2. On cumulative full-application reconstruction workflows, GPT-6 Astra achieved 49.2 percent success while Claude Opus 5 reached 28.8 percent.

    GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows
  3. The paper was submitted to arXiv on September 16, 2026.

    Submitted on 16 Sep 2026

Nothing appears on this site that we have not opened and linked.

Filed under