Models tested against my personal workflows, rather than an arbitrary benchmark.
Click a row for its task scores, a column header to sort.
Eight questions across six repositories, answered with exact paths and line numbers.
| Model | Score | Cost | Time | Tokens | Run |
|---|
Compress a 3,853-line business-rule spec and a 723-line client audit, in under 40 lines each.
| Model | Score | Cost | Time | Tokens | Run |
|---|
Name every file a known change has to touch. The real fixes touched 25 and 19 files.
| Model | Score | Cost | Time | Tokens | Run |
|---|
Ship a CMS section end to end, and apply a design batch that then gets corrected mid-task.
| Model | Score | Cost | Time | Tokens | Run |
|---|
Three storefront-header and cart/menu diffs: the header change with two injected requirement violations, the same header change authentic, and a small correct mobile fix.
| Model | Score | Cost | Time | Tokens | Run |
|---|