Claude Code's /doctor prompt-audit, tested on nine planted problems — Opus and Sonnet caught all nine, Haiku missed the broken commands
The audit runs on your session's model. On our test repository, Opus and Sonnet flagged all nine planted problems in every run, in about a minute. Haiku missed the nonexistent test command and the missing deploy script every time, and in a preliminary run said the test command was fine.
Update — October 7, 2026: Haiku 5.5 changes the Haiku result. Anthropic released Claude Haiku 5.5 today, and Claude Code 2.1.293 made it the default for the
haikualias on the Anthropic API. Rerun on the same fixture, Haiku 5.5 found all nine planted problems in three of three runs and listed no controls, for about $0.04 a run; Haiku 4.5, rerun the same day, found three, five and eight. The Haiku 4.5 results below stand as a snapshot of that model, which stays available by its full ID. The rerun, and the price step at 100K tokens that one Haiku 5.5 run crossed, are in Haiku 5.5’s 90% price cut stops at 100K tokens.
The 30-second version
On 25 September, Claude Code 2.1.283 added /doctor prompt-audit. It has Claude review your CLAUDE.md, AGENTS.md, skills and subagents for three kinds of problem: instructions written for older models, references to files or commands that don’t exist, and files that contradict each other. It reports and proposes edits; it changes nothing until you ask.
We built a small repository with nine planted problems and six instructions kept as controls, then ran the audit three times on each of Opus 5.5, Sonnet 5.5 and Haiku 4.5:
- Opus 5.5: all nine problems in every run, and none of the controls listed as a problem. About $0.75 and a minute per run, at API prices.
- Sonnet 5.5: all nine in every run, plus one low-confidence finding on a control. About $0.41.
- Haiku 4.5: five, seven and six. It missed a nonexistent npm script and a missing deploy script in all three runs; a separate preliminary run called the script fine. It also wanted to replace the documented
ultrathinkkeyword with effort settings, every time. - Git reported no changes in the repository after any of the nine runs.
Verdict: worth running on Opus or Sonnet, especially after you change models or build commands — then check every proposed change. The audit runs on your session’s model, so switch with /model first if you normally run Haiku.
What the audit does
/doctor prompt-audit (also /checkup prompt-audit) needs Claude Code 2.1.283 or later. It runs through the bundled claude-api skill, so it is unavailable if you turn that skill off.
By default it covers your CLAUDE.md, CLAUDE.local.md and AGENTS.md files, plus the rules, skills, commands, subagents and output styles under .claude/ and ~/.claude/. That default includes your personal configuration. To keep it to one repository or one file, pass a path: /doctor prompt-audit ., or /doctor prompt-audit .claude/skills/deploy.
Part of what it checks matches the advice Anthropic gave with Opus 5.5. Addy Osmani’s guide for the model, published 22 September, says: “Delete ‘think carefully’ lines. Opus 5.5 already thinks before every reply.” Our delete-your-CLAUDE.md piece covers the research on what trimming instruction files does and doesn’t change.
How we tested
The test repository, acme-api, has a package.json with three scripts (test, lint, build), two source files indented with two spaces, a CLAUDE.md, an AGENTS.md, one skill and one subagent. We planted nine problems and kept six correct instructions as controls:
| # | Planted problem | Where |
|---|---|---|
| 1 | ”think step by step and think carefully before every single response” | CLAUDE.md |
| 2 | All-caps pressure: “NEVER EVER skip this. THIS IS CRITICAL!!!” | CLAUDE.md |
| 3 | npm run test:unit, a script that doesn’t exist | CLAUDE.md |
| 4 | ./scripts/deploy.sh staging, a file that doesn’t exist | CLAUDE.md |
| 5 | ”Indent with tabs” against “Indent with 2 spaces” | CLAUDE.md vs AGENTS.md |
| 6 | @docs/architecture.md, an import of a missing file | CLAUDE.md |
| 7 | ”Use XML tags like <thinking></thinking> to reason before you answer” | CLAUDE.md |
| 8 | ”As a large language model, you may not be able to run commands” | skill |
| 9 | model: claude-3-5-sonnet-20241022, a retired model | subagent |
The six controls were instructions that match the repository — npm install, npm run lint, npm test, npm run build with git tag, and a rule that only src/db/client.ts may import pg — plus ultrathink when a change touches billing code. Anthropic’s 2.1.283 release notes say the audit keeps thinking keywords Claude Code documents, so we counted a proposal to replace ultrathink against it.
Each run was claude -p "/doctor prompt-audit ." — a non-interactive run — on Claude Code 2.1.289, in a fresh copy of the repository, with project settings only, three at a time. Nine runs, $4.08 in total at API list prices, 529 seconds of summed run time, on 4 October. We read every report. A planted problem counted as caught when the report flagged its line. A control counted against the model when the report listed it as a finding; a line the report kept with an optional suggestion did not count.
Results
| Planted problem | Opus 5.5 | Sonnet 5.5 | Haiku 4.5 |
|---|---|---|---|
| 1. “think step by step” | 3/3 | 3/3 | 3/3 |
| 2. All-caps pressure | 3/3 | 3/3 | 3/3 |
3. Nonexistent npm run test:unit | 3/3 | 3/3 | 0/3 |
4. Missing scripts/deploy.sh | 3/3 | 3/3 | 0/3 |
| 5. Tabs vs 2 spaces | 3/3 | 3/3 | 3/3 |
6. Missing @docs/architecture.md | 3/3 | 3/3 | 1/3 |
7. <thinking> tags | 3/3 | 3/3 | 2/3 |
| 8. “As a large language model…“ | 3/3 | 3/3 | 3/3 |
| 9. Retired model ID | 3/3 | 3/3 | 3/3 |
| Found per run | 9, 9, 9 | 9, 9, 9 | 5, 7, 6 |
| Controls listed as findings | 0 | 1 (low confidence) | 3 (ultrathink, every run) |
Opus and Sonnet checked the repository, not just the text. They read package.json to show test:unit doesn’t exist and that AGENTS.md already says npm test, confirmed there is no scripts/ folder, and checked that only src/db/client.ts imports pg before keeping that rule. On the indentation conflict, Opus in every run and Sonnet in two of three pointed out that the code uses two spaces, and both left the choice to us. Findings came with proposed edits.
Sonnet’s one finding on a control, in one run, was a low-confidence suggestion that the git tag vX.Y.Z step should say where the version number comes from — a reasonable point, counted by our rule. Its reports also offered optional clarifications on lines they kept.
What Haiku got wrong
It missed the commands that don’t exist. A nonexistent npm script and a missing deploy script belong to the “references to files or commands that don’t exist” the documentation says the audit looks for. Haiku missed both in all three runs. In a separate preliminary run, it looked at the test:unit line and called it “operational, not model-related, and remains current”.
It wanted to replace ultrathink with effort settings. In every run it flagged the keyword as outdated and proposed output_config: {effort: "xhigh"} or effort: "max". Those are request parameters for the Claude API; writing them in a CLAUDE.md doesn’t set the API’s effort. Opus and Sonnet kept the line in every run. (Whether ultrathink does anything inside a CLAUDE.md is a separate question, and we didn’t test it: the documentation describes it as a keyword you put in your prompt.)
It sided against the code on the contradiction. All three Haiku runs proposed changing AGENTS.md to tabs, though both source files use two spaces; the third also called the choice a project decision. In one run it swapped the “think step by step” block for another sentence telling the model to reason carefully.
Cost and time
| Model | Mean cost per run, at API prices | Mean time per run | Turns reported |
|---|---|---|---|
| Opus 5.5 | $0.75 ($0.74–$0.77) | 62 s | 10–13 |
| Sonnet 5.5 | $0.41 ($0.39–$0.42) | 45 s | 12 |
| Haiku 4.5 | $0.20 ($0.17–$0.23) | 69 s | 10–14 |
Haiku was the cheapest and the slowest. Turns are the counts Claude Code reported for each run. On a Pro or Max plan, runs within your allowance use included usage, and usage beyond it is billed as usage credits. Our repository is tiny; larger configurations may cost more, but we didn’t measure how cost scales.
We ran it on our own CLAUDE.md
We pointed it at this site’s own CLAUDE.md, on Opus 5.5: 13 turns, 108 seconds and $1.05 at API prices. It read the file, then the layouts, config and scripts it needed to check it, and listed nine findings. We confirmed three as real problems:
- A rule that sends you to the wrong file. Our table-width rule said the CSS lives in
tailwind.config.mjs. It is in a layout file, on a wrapper a build plugin adds. - A contradiction inside one file. Our plain-language rule bans one Chinese term as jargon. Our Chinese style list, a few lines later, recommends the same term.
- Script names without their folders:
gen-og.pyandgen-llms.pylive underscripts/.
The rest were judgment calls: a status note we had added the week before that will go stale, section headings it found too emphatic, and a rule duplicated in our AGENTS.md that has drifted. One statement blurred two things: it said both AGENTS.md and CLAUDE.md are loaded in our repository, after reading both itself. By default Claude Code loads only CLAUDE.md at startup when both exist, as we measured last week. It couldn’t run git blame either, because a non-interactive run can’t approve the command. We fixed the CSS location and the script paths; the wording contradiction is still open.
How to use it
- Run it on Opus or Sonnet. The audit uses your session’s model. If your default is Haiku, run
/model sonnetfirst. (Aliases follow the newest model; we tested the IDs listed below.) - Pass a path to choose which instruction files it audits.
/doctor prompt-audit .keeps it to the repository you’re in; without a path it also covers~/.claude/. It still reads other project files to check facts. - Decide contradictions yourself. It flags them and may point out which side the code supports. The choice stays with you.
- Run it again after you switch models or change build commands. Those are the changes that turn correct instructions into outdated ones.
- Read the proposed edits before applying them. Nothing changes until you ask.
What we didn’t test
- Large, real configurations. Our repository is small by design, so every answer is known. Beyond it we audited only our own
CLAUDE.md, once. - Fable 5.1.
- Whether acting on the findings improves results. The audit suggests what may be outdated, not whether removing it makes the agent better; our delete-your-CLAUDE.md piece covers what the research found.
- More than three runs per model. Opus and Sonnet flagged all nine planted problems in every run, though their explanations and extra suggestions varied; Haiku’s totals varied.
Companion reading
- Anthropic deleted 80% of Claude Code’s system prompt — the research on what instruction files do and don’t change.
- How big can your instructions file get before it hurts? — the size side of the same question.
- How to write a great CLAUDE.md — what belongs in the file in the first place.
- Claude model tiers and effort levels — what
ultrathinkand effort actually do. - Claude Code reads AGENTS.md now — which instruction files load, and when.
Sources
- Anthropic — Manage Claude’s memory: audit your instruction files. What the audit looks for, its default scope, the path argument, and that nothing changes until you ask.
- Anthropic — Commands.
/doctor [prompt-audit [path]], the/checkupalias, and the 2.1.283 requirement. - Anthropic — Claude Code changelog. 2.1.283: the audit added, and “thinking keywords that Claude Code documents are kept”.
- Anthropic — Model configuration: use ultrathink for one-off deep reasoning.
ultrathinkas a keyword in your prompt. - Addy Osmani, Anthropic — Getting the most out of Opus 5.5 in Claude and Claude Code, 22 September 2026.
- npm — @anthropic-ai/claude-code. 2.1.283 published 25 September.
Our own measurements: nine claude -p runs of /doctor prompt-audit . on macOS on 4 October 2026, Claude Code 2.1.289, three each on claude-opus-5-5, claude-sonnet-5-5 and claude-haiku-4-5-20251001, with project settings only, each in a fresh copy of the test repository. Costs are the totals Claude Code reported at API list prices; $4.08 for the nine runs, plus $1.05 for one Opus run on our own CLAUDE.md. Every report was read and scored by hand against the nine planted problems and six controls.
FAQ
What does /doctor prompt-audit do? It has Claude check your instruction files for instructions written for older models, references to things that don’t exist, and contradictions between files, then proposes edits. Added in 2.1.283; it changes nothing until you ask.
Does the model matter? On our test, yes. Opus 5.5 and Sonnet 5.5 flagged all nine planted problems in every run. Haiku 4.5 flagged five to seven and missed the nonexistent test command and deploy script every time.
What does it cost? About $0.75 per run on Opus, $0.41 on Sonnet and $0.20 on Haiku, at API prices, on our small repository. On a subscription it uses included usage first.
Will it change my files? Not unless you ask. Git reported no changes after any of our nine runs.
Should I delete “think step by step”? Every model flagged it in every run, and Anthropic’s Opus 5.5 guide says to delete “think carefully” lines. We didn’t test whether removing it changes results. Keep what the model can’t work out for itself, such as your build and test commands.
Was this helpful?