notis.ai

What actually shipped. Every claim carries the source it rests on.

Research

ProgramDistill, a benchmark built from 26 web apps, scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8.

2026-09-28 · that day's edition

Built from 26 functioning reference applications, the benchmark scores GPT-6 Astra at 49.2 percent and Claude Opus 5 at 28.8 percent on full-app reconstruction.

What this rests on

  1. The benchmark draws on 26 fully functional reference web applications, from which the researchers extracted 1,975 replay-verified behaviors and built 4,063 tasks.

    ProgramDistill: benchmarking coding agents on feature discovery in real applications · arXiv · 2026-09-16
    GPT-6 Astra achieved 49.2% success while Claude Opus 5 reached 28.8% on cumulative full-application reconstruction workflows.

  2. On cumulative full-application reconstruction workflows, GPT-6 Astra achieved 49.2 percent success while Claude Opus 5 reached 28.8 percent.

    ProgramDistill: benchmarking coding agents on feature discovery in real applications · arXiv · 2026-09-16
    GPT-6 Astra achieved 49.2% success while Claude Opus 5 reached 28.8% on cumulative full-application reconstruction workflows.

  3. The paper was submitted to arXiv on September 16, 2026.

    ProgramDistill: benchmarking coding agents on feature discovery in real applications · arXiv · 2026-09-16
    GPT-6 Astra achieved 49.2% success while Claude Opus 5 reached 28.8% on cumulative full-application reconstruction workflows.

We checked every sentence above against its source by opening it. Nothing appears on this site that we have not opened and linked.

Filed under

Also that day