Which AI assistant gets it done, and how fast?
Claude Opus 5.5 passes 80% of everyday tasks; the best app, Pally, passes 61%. Every run is graded from the record of what the assistant actually did, and published in full.
Leaderboard
80% 69%–88% | 100% | 20s | 2 | 96.8% | 85.7% | 100% | 98.3% | |
76% 65%–85% | 100% | 14s | 0 | 100% | 84.6% | 100% | 100% | |
72% 60%–81% | 98.5% | 30s | 2 | 96.9% | 85.7% | 100% | 100% | |
67% 55%–77% | 100% | 17s | 3 | 96.9% | 38.4% | 100% | 100% | |
65% 53%–76% | 100% | 14s | 3 | 95.2% | 50% | 100% | 98.3% | |
61% 49%–72% | 84.8% | 42s | 2 | 100% | 61.5% | 100% | 100% | |
7Muse | 58% 46%–69% | 85.7% | 1m 37s | 1 | 100% | 71.4% | 100% | 100% |
56% 44%–67% | 95.5% | 26s | 3 | 96.9% | 38.4% | 100% | 100% | |
51% 39%–63% | 95% | 56s | 0 | 100% | 40% | 100% | 100% |
What we found
- The best bare model outscored every appClaude Opus 5.5 passed 80% of the everyday tasks on its own; the best app, Pally, passed 61%. Their 95% intervals overlap, so the gap may not hold up.
- Most failures are unfinished jobs188 of 207 failed runs missed something the task called for: the wrong action, a wrong end state, or a wrong answer.
- Assistants guess instead of askingWhen a task hinged on something only the user knew, the assistant guessed in 41 of 109 runs (38%).
- The apps claim things they didn’t doThe apps said they had done something the services’ logs don’t show in 22 of 189 runs; bare models in 4 of 394.
- Bare models act before you say yesBare models changed, sent or bought something before the user said yes in 11 of 386 runs; the apps in 0 of 176.
- No assistant acted on made-up detailsAcross 593 runs, no assistant put an ID, date, amount or recipient into an action that it hadn’t got from the user or a real result.
Results by task: how each assistant did on all 34 tasks
One square per run: PassFailUnsafe fail Hardest task first; pick a square for that run in full.