Benchmarks
Distill with GPT-6.1 Sol auto and GPT-6 Luna auto used 33.8% less time than Codex with GPT-6.1 Sol high on three small Python tasks, each repeated three times. Both agents passed all nine runs and independent review of their actual diffs.
| Agent | Total time | Time saved | Estimated cost (USD) | Cost saved | Passing runs |
|---|---|---|---|---|---|
| Distill Sol auto + Luna auto | 308.18 s | 33.8% | $0.19 | 67.2% | 9/9 |
Codex Sol high --yolo
| 465.27 s | 0% (baseline) | $0.59 | 0% (baseline) | 9/9 |
Dollar estimates use Standard credit rates and $0.04 per credit, based on OpenAI's published equivalence of 2,500 credits to $100. Savings compare unrounded totals against Codex. Excluding Codex's zero-output requests from its cost estimate gives a 52.6% reduction. Every call remains in the usage totals, including workers, titles and retries. Both agents used the same ChatGPT account.
These results cover the tested tasks. Dollar figures are reference estimates; actual credit prices can vary by plan. They do not measure the included subscription allowance or reduce the fixed monthly price. BENCHMARKS.md includes the protocol, raw results, rejected attempts and total usage audit.
Changes
- Preserve ChatGPT OAuth when resolving child models and auxiliary calls.
- Skip dashboard summary calls when a session runs without an interactive attachment.
- Preserve native Codex session headers and optional tool arguments.
- Keep Worker instructions through child startup and batch local reads and checks.
- Supply the actual repository diff and recorded check results to Main for review of fresh foreground Workers in headless sessions. Incomplete evidence still requires verification.
- Preserve each function's existing module and reject duplicate implementations.
- Add the comparison table and benchmark report link to the README.
Verification
The final cohort passed all 18 executions, independent behavior and scope checks, function-ownership checks and diff review. Accounting is complete for all 119 model calls. The focused checks passed 18 Python tests and 22 Rust tests.