ainotis Join
My notis

Checked fact 16928 Sep 2026Research

On cumulative full-application reconstruction workflows, GPT-6 Astra achieved 49.2 percent success while Claude Opus 5 reached 28.8 percent.

The exact words it rests on

GPT-6 Astra and Claude Opus 5 achieve 49.2% and 28.8% success on cumulative workflows

What the source said when we opened it, on 28 Sep 2026.

The source

ProgramDistill: benchmarking coding agents on feature discovery in real applications
arXiv · 2026-09-16

Checked

Checked by the notis newsroom on , against the source above.

The number in it

  • 49.2% · GPT-6 Astra, benchmark score (ProgramDistill cumulative workflows)
  • 28.8% · Claude Opus 5, benchmark score (ProgramDistill cumulative workflows)

In the story

ProgramDistill, a benchmark built from 26 web apps, scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8. 28 Sep 2026

Cite this fact

Anyone may quote this address. It does not change; if we correct the story, this page says so.