Revise Blog - Building a post-AI word processor
The best AI model for proofreading (we tested them all)
We’ve tested 47 AI models on a large scale proofreading benchmark. Here's how they stack up.
You probably use AI to review your writing, but have you ever wondered which models you should be using for this? We did, so we built a proofreading benchmark to answer that with hard data.
ErrataBench uses a dataset of English text totaling 99,000 words - literature, legal writing, technical manuals - corrupted with a wide variety of realistic writing errors. Each model is asked to find and fix as many of these as it can, and judged by its completeness. Here is what we learned.
TL;DR: which model is the best?
As of September 22, 2026, GPT-5.6 Sol tops the benchmark at 96.5% with medium reasoning. But the top six models all land within two points of each other, while their cost per benchmark run ranges from $0.41 to $1.88. For most proofreading, the last point of accuracy isn’t worth paying three times as much for, so our recommendations lead with value:
- Best value overall: GPT-6 Sol. With no reasoning it scores 95.4% for $0.41 per run - about one point behind GPT-5.6 Sol at a third of the price, fixing 109 errors per dollar to Sol’s 37. It’s also fast, at 38 fixes per minute.
- Best Claude: Opus 5.5. With high reasoning it scores 94.7% for $0.46 per run - one point behind Claude Opus 5 (95.7%) at half the price, or about twice as many fixes per dollar. It’s also the most careful model we tested: its 0.4% bad-fix rate is the lowest of any model scoring over 90%. The trade-off is completeness - it leaves 4.9% of errors untouched, more than the other flagships.
- Only if every point matters: GPT-6 Astra. It scores 96.1% with low reasoning, within half a point of GPT-5.6 Sol, for $1.13 per run - and it fixes errors nearly five times as fast (40.3 fixes per minute to Sol’s 8.4). GPT-5.6 Sol itself is the most accurate model, missing just 2.3% of errors, but costs more than Astra and runs about five times slower.
Some other noteworthy mentions based on the full data:
- Opus 5 likes its reasoning dial low; Opus 5.5 does not. Opus 5 scores best with low reasoning (its no-reasoning variant drops all the way to 86.5%), and the older Opus models (4.6, 4.7, 4.8) do best with no reasoning at all. Opus 5.5 reverses the pattern: it scores highest with high reasoning and drops to 91.6% on low.
- Gemini 3.7 Flash is the new budget pass - it scores 91.4% for about $0.12 per run while fixing 31.4 errors per minute.
- GLM 5.3 Flash is the cheapest model to crack 90%, scoring 90.1% for about $0.07 per run.
- Grok 4.5 remains xAI’s best at 93.4%, but GPT-6 Sol now beats it on quality, price, and speed. The newer Grok 4.6 scores slightly lower (92.9%) at a higher price and runs at about 4 fixes per minute.
- There are certain models it just never makes sense to use, such as Claude Sonnet (inferior to Gemini and GLM) and Grok 4.2 - beaten on both quality and price by xAI’s own 4.3 and 4.5.
The top 6 models by quality overall are below. See the full results for more data and scatterplots.
| # | Model | Maker | Reasoning | Fix rate | Fixes per $1 | Fixes per min |
|---|---|---|---|---|---|---|
| 1 | GPT-5.6 Sol | OpenAI | Medium | 96.5% | 37.4 | 8.4 |
| 2 | GPT-6 Astra | OpenAI | Low | 96.1% | 41.8 | 40.3 |
| 3 | Claude Opus 5 | Anthropic | Low | 95.7% | 44.3 | 22.5 |
| 4 | GPT-6 Sol | OpenAI | — | 95.4% | 109.0 | 38.0 |
| 5 | GPT-5.5 | OpenAI | High | 94.8% | 25.5 | 5.2 |
| 6 | Claude Opus 5.5 | Anthropic | High | 94.7% | 95.2 | 30.6 |
If you want to try different models, you can! Revise is an agentic document editor with support for GPT, Claude, Gemini, and Grok all in one.
Try it yourself. Paste or upload the document you want to proofread below and edit it with AI right in your browser - try it with GPT-6 Luna and Claude Haiku for free, or access all the top models with a paid plan. No installation or API key required.
How we tested every model
A proofreading benchmark only means something if the test is realistic. Here’s the short version of how ErrataBench works:
- Real text, real errors. We start from genuine source documents across 11 datasets - novels, legal text, scientific writing, nearly 99,000 words in all - and use a model pair to seed small, realistic mistakes drawn from a taxonomy of error categories, with a second model reviewing each one before it’s accepted.
- No cheat sheet. Each model is asked to proofread carefully, but is never told what kinds of errors are hidden or how many there are - just like a real editing job.
- The same tools for everyone. Models work through a simple agent loop with find-and-replace tools, get the same text chunks and the same number of turns, and are graded by an LLM judge on whether each fix is actually correct.
Every model gets the identical prompt, tools, and text - so the only variable is the model itself. Full methodology, source code, and raw results are published and reproducible.
Beyond the headline fix rate, two numbers tracked in the full results matter just as much: missed errors - the ones a model never touched (a completeness problem) - and bad fixes, the changes it attempted that were actually wrong. The second is the dangerous one - a model that “corrects” text incorrectly makes your writing worse. The best models keep bad fixes near 1%.
Accuracy vs. cost: the smart-money picks
The most accurate model isn’t always the one you should reach for. Proofreading is often high-volume - a long report, a whole manuscript, a stack of documents - so cost and speed matter.
At the very top, GPT-5.6 Sol leads on quality at about $1.27 per benchmark run, and GPT-6 Astra comes within half a point for $1.13. Those last points are expensive, though. GPT-6 Sol reaches 95.4% for about $0.41 with no reasoning - cheaper than every model that beats it, and far cheaper than GPT-5.5 ($1.88), which scores lower. Claude Opus 5.5 makes the same trade on the Claude side: 94.7% for about $0.46, against Opus 5’s 95.7% for $0.96. Between the two, GPT-6 Sol edges Opus 5.5 on quality, price, and speed, while Opus 5.5 makes fewer bad fixes.
Below the flagships, Gemini 3.1 Pro reaches 92.1% for roughly $0.37, Gemini 3.7 Flash reaches 91.4% for about $0.12, and GPT-6 Luna reaches 90.9% for about $0.15. Kimi K3 matches Gemini 3.1 Pro’s 92.1%, but at about $0.91 per run and a much slower pace, Gemini is the better buy at that quality level.
That’s the whole argument for not marrying a single model: use a fast, cheap one for quick passes and escalate to a flagship for the writing that really matters. You can explore the full accuracy-vs-cost tradeoff - plus speed and consistency - on the interactive ErrataBench scatterplot.
What even the best models still get wrong
No model is perfect, and the errors they miss aren’t random. Aggregated across every model we tested, the easiest category to catch is spelling & word form - the obvious typos a basic checker would flag too, resolved about 92% of the time. The hardest is meaning, flow & style, resolved only around 81% of the time - the subtle, judgment-heavy issues where reasonable editors can disagree. The full results break performance down by error category.
The practical lesson: AI proofreading is excellent, but it’s a collaborator, not an oracle. You still want to see and approve what it changes - which is why the how of AI proofreading matters as much as which model.
The best way to proofread with AI (not just the best model)
Picking a model is half the battle. The other half is the tool you run it in. Pasting paragraphs into a chatbot and copying results back is slow, loses your formatting, and - worst of all - gives you no easy way to see exactly what changed.
Revise is a word processor with the AI built into the editor, designed for exactly this:
- Every top model, one click away. Because Revise is multi-provider, you can proofread with GPT-6 Astra, run a cheaper pass with GPT-6 Sol or Claude Opus 5.5, and get a second opinion from Gemini 3.1 Pro - all in the same document, no accounts or API keys to juggle.
- Whole-document context. The AI reads your entire draft, so it catches inconsistencies a paragraph-at-a-time chatbot never would - a term spelled two ways, a tense that drifts.
- Tracked changes you control. Every fix appears as a red/green diff you accept or reject one by one. The AI proposes; you decide. Nothing is overwritten silently.
- Full revision history. Scrub back through every version to see exactly how the document evolved and which edits came from you versus the AI.
In other words, Revise turns “which model is best” from a one-time bet into a dial you can turn per task - with the review controls that make it safe to let AI touch a document that matters.
The bottom line
What’s the best AI model for proofreading?
For most people, GPT-6 Sol with no reasoning: it scores 95.4%, about one point behind the leader, for a third of the price. If you prefer Claude, Claude Opus 5.5 with high reasoning scores 94.7% at half the price of Opus 5 and makes the fewest bad fixes of any top model. The most accurate model overall is GPT-5.6 Sol at 96.5%, with GPT-6 Astra close behind at 96.1% and nearly five times faster.
Is ChatGPT or Claude better at it?
ChatGPT leads: GPT models hold first, second, and fourth place, and GPT-5.6 Sol beats Claude Opus 5 by 0.8 percentage points. GPT-6 Astra also outscores Opus 5 while fixing 40.3 errors per minute to its 22.5, though Opus 5 is a little cheaper per run ($0.96 vs. $1.13). Claude’s edge is caution: Claude Opus 5.5 makes bad fixes on just 0.4% of errors, the lowest of any model scoring over 90%.
How accurate is AI proofreading, really?
The leaders find and correctly fix more than 90% of errors (the best, GPT-5.6 Sol, hits 96.5%) and rarely break anything - but they’re strongest on mechanical mistakes and weakest on judgment calls involving meaning, flow, and style. Treat the AI as a collaborator and review its changes.
And since the models keep leaping ahead of each other, the smartest move isn’t to pick one forever - it’s to use a tool that lets you switch. Try it on your own writing: paste a document and proofread it with the top AI models. It’s free to use with basic AI, and with a paid upgrade it offers the most powerful AI features of any word processor. And if you want to go deeper on the data, the full ErrataBench results are open and interactive.
Quoting this post? All the data here is free to reuse with a link back.


