Research
ProgramDistill, a benchmark built from 26 web apps, scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8.
2026-09-28 · that day's edition
Built from 26 functioning reference applications, the benchmark scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8 percent on full-app reconstruction.
What this rests on
-
The benchmark draws on 26 fully functional reference web applications, from which the researchers extracted 1,975 replay-verified behaviors and built 4,063 tasks.
ProgramDistill: benchmarking coding agents on feature discovery in real applications · arXiv · 2026-09-16
GPT-6 Astra achieved 49.2% success while Claude Opus 5 reached 28.8% on cumulative full-application reconstruction workflows.
-
On cumulative full-application reconstruction workflows, GPT-6 Astra achieved 49.2 percent success while Claude Opus 5 reached 28.8 percent.
ProgramDistill: benchmarking coding agents on feature discovery in real applications · arXiv · 2026-09-16
GPT-6 Astra achieved 49.2% success while Claude Opus 5 reached 28.8% on cumulative full-application reconstruction workflows.
-
The paper was submitted to arXiv on September 16, 2026.
ProgramDistill: benchmarking coding agents on feature discovery in real applications · arXiv · 2026-09-16
GPT-6 Astra achieved 49.2% success while Claude Opus 5 reached 28.8% on cumulative full-application reconstruction workflows.
We checked every sentence above against its source by opening it. Nothing appears
on this site that we have not opened and linked.
Filed under
Also that day