How We Test

Every verdict on this site traces back to hands-on use, and this page explains the machinery behind that: what we test with, how comparisons are actually run, how facts get checked, and what to do when we get something wrong. If a review or comparison here ever seems to contradict this page, this page wins — tell us and we’ll fix the article.

What We Test

Our coverage spans three surfaces. First, the chat products: the claude.ai web app, the desktop and mobile apps, and — for comparisons — the equivalent products from OpenAI, Google, and others. Second, Claude Code, Anthropic’s agentic coding tool, in the terminal and IDE. Third, the API, for developer-facing articles about pricing, rate limits, and integration. When an article’s conclusions only apply to one of these surfaces, the article says so.

Our Testbed

As of July 2026, the current Claude models we test against are Fable 5, Opus 4.8, Sonnet 5, and Haiku 4.5. Unless an article states otherwise, we use each product at its default settings — no custom system prompts, no unusual configuration — because that’s how most readers will experience it. Our access runs through a paid claude.ai Pro plan and a separately billed API account, both purchased at the same rates anyone else pays. Nothing we test on is provided free by a vendor.

How Comparisons Are Run

When we compare tools, each one gets the same set of real-world tasks — the drafting, debugging, research, and document work we already do daily — rather than synthetic puzzles built to flatter a particular model. The judgments that come out of that are qualitative, and we label them as what they are: author assessment, formed at the keyboard. We do not publish invented benchmark numbers. If you see a figure like an accuracy score or a leaderboard position here, it comes from a named third party and carries a date, because a benchmark result without its source and vintage is closer to decoration than evidence.

How Facts Are Verified

Factual claims — prices, model names, context windows, feature availability — are checked against official Anthropic documentation and announcements before anything else, and secondary sources only fill gaps the official material leaves. Claims are dated in the text so you can judge their freshness yourself. For fast-moving facts about competitors, where a price or feature can change between our publish date and your visit, we hedge explicitly with “as of” phrasing rather than presenting a snapshot as a permanent truth.

Corrections

When we get something wrong, we want to hear about it. Email [email protected] with the article URL and the issue; verified corrections jump our queue. Fixes are made in the article itself and reflected in that article’s Updated date, so the page you’re reading always shows when it last changed. We don’t quietly rewrite history — if a recommendation was wrong, the revised article says what changed.

Honest Limitations

Three caveats, stated plainly. We are a small editorial team, not a benchmarking lab — our strength is depth of daily use, not breadth of controlled trials. AI products change weekly: a limit, price, or behavior we describe accurately today can be different next month, which is why every article carries its dates. And on any point where this site and Anthropic’s official documentation disagree, the official docs prevail — treat us as the experienced colleague, and them as the source of record.

Questions about anything on this page? Get in touch — methodology feedback is some of the most useful mail we get.