# AssistantEval > AssistantEval is a verifiable benchmark for personal AI assistants. Every run is graded from the record of what the assistant actually did, and every run is published in full. Results updated Oct 5, 2026. The leaderboard is licensed CC BY 4.0: cite “AssistantEval (assistanteval.com)”. Each run gives an assistant a small, realistic world (email, calendar, contacts, shopping and more) and a simulated user who knows things the assistant has to ask for. The grader reads the services' own logs and checks that the assistant asked before acting, asked for what it needed, used only real information, was truthful about what it did, told the user what it had done, and got the job done. ## What we found - The best bare model outscored every app. Claude Opus 5.5 passed 80% of the everyday tasks on its own; the best app, Pally, passed 61%. Their 95% intervals overlap, so the gap may not hold up. - Most failures are unfinished jobs. 188 of 207 failed runs missed something the task called for: the wrong action, a wrong end state, or a wrong answer. - Assistants guess instead of asking. When a task hinged on something only the user knew, the assistant guessed in 41 of 109 runs (38%). - The apps claim things they didn’t do. The apps said they had done something the services’ logs don’t show in 22 of 189 runs; bare models in 4 of 394. - Bare models act before you say yes. Bare models changed, sent or bought something before the user said yes in 11 of 386 runs; the apps in 0 of 176. - No assistant acted on made-up details. Across 593 runs, no assistant put an ID, date, amount or recipient into an action that it hadn’t got from the user or a real result. ## Leaderboard: assistant products Consumer assistants, used the way a person would use them, on short everyday tasks. Ranked by the share of tasks passed outright. 1. [Pally](https://assistanteval.com/assistants/pally): Pally passed 40 of 66 everyday tasks (61%). 2. [Muse](https://assistanteval.com/assistants/muse): Muse passed 37 of 64 everyday tasks (58%). 3. [Grok Bot](https://assistanteval.com/assistants/grok_bot): Grok Bot passed 32 of 63 everyday tasks (51%). ## Leaderboard: frontier models Bare models given the same services as tools and a neutral prompt, on the same everyday tasks as the products. Ranked by the share of tasks passed outright. 1. [Claude Opus 5.5](https://assistanteval.com/assistants/claude-opus-5.5): Claude Opus 5.5 passed 52 of 65 everyday tasks (80%). 2. [GPT-6 Astra](https://assistanteval.com/assistants/gpt-6-astra): GPT-6 Astra passed 51 of 67 everyday tasks (76%). 3. [Claude Fable 5.1](https://assistanteval.com/assistants/claude-fable-5.1): Claude Fable 5.1 passed 49 of 68 everyday tasks (72%). 4. [Gemini 3.1 Pro](https://assistanteval.com/assistants/gemini-3.1-pro-preview): Gemini 3.1 Pro passed 44 of 66 everyday tasks (67%). 5. [Grok 4.7](https://assistanteval.com/assistants/grok-4.7): Grok 4.7 passed 43 of 66 everyday tasks (65%). 6. [Muse Spark 1.3](https://assistanteval.com/assistants/muse-spark-1.3): Muse Spark 1.3 passed 38 of 68 everyday tasks (56%). ## Read more - [Method](https://assistanteval.com/method): how tasks are built and runs are graded - [Leaderboard as CSV](https://assistanteval.com/data/results.csv): one row per assistant (CC BY 4.0) - [Data terms](https://assistanteval.com/terms): the run records under https://assistanteval.com/data/ (each run's conversation, actions and verdict) may be read, verified and cited, but not used to train AI models ## Optional - Contact: hello@assistanteval.com