Your Opus 5 system prompt has a max_tokens cap that worked fine for months. Switch to Opus 5.5 without re-checking it, and that same cap can now silently truncate responses, because hidden thinking tokens eat into the budget you thought was reserved for output. That's just one of several places where a straight swap from Opus 5 to Opus 5.5 can quietly break things.
Anthropic published prompting guidance for Opus 5.5 on the model's September 22 launch. It's unusually specific about what to re-test rather than assume (Search Engine Journal). The guide says existing Opus 5 prompts should still work without edits, and that Opus 5 guidance remains a reasonable starting point.
But "still works" and "still optimal" are different claims. The gap between them is where cost and quality get wasted. This checklist turns Anthropic's guidance into a step-by-step audit you can run before you flip production traffic to Opus 5.5.
One caveat up front: everything below synthesizes Anthropic's own published guidance and reported internal test results. It is not independent benchmarking. Treat the verification steps here as your way of confirming Anthropic's claims hold for your specific prompts and traffic, not as a substitute for that confirmation.
Step 1: Find Every Effort Setting In Your Codebase
Start by grepping your codebase and prompt configs for anywhere you set an effort level, whether that's hardcoded, pulled from an environment variable, or baked into a request template. Opus 5 defaulted to high effort. Opus 5.5 defaults to medium, one level down (Search Engine Journal). If your app never explicitly sets effort, it now runs at medium on Opus 5.5 with zero code changes — which may be exactly what you want, or a silent quality shift you didn't plan for.
List every place effort is set and what value it's set to. If you copied a "high" setting from Opus 5 into Opus 5.5 config to preserve behavior, flag that line for retesting in Step 5. Anthropic's guidance is explicit that this is the first thing to re-check, not the last: "Don't assume your Claude Opus 5 setting is still the right one for Opus 5.5" (Search Engine Journal).
Step 2: Audit Think-Carefully Lines In Chat Prompts
Search your system prompts for phrasing like "think carefully before responding," "take your time," or similar instructions meant to slow the model down for multistep reasoning. These lines were a common workaround under Opus 5. Anthropic's own effort documentation for Claude Opus 4.7 recommended exactly this kind of line when effort had to stay low for speed (Search Engine Journal).
Opus 5.5 changes the mechanism. Thinking can't be turned off on this model, and Opus 5.5 decides for itself how much to think, with effort as the primary lever instead of prompt phrasing. In Anthropic's own chat-product test, removing a think-carefully line made replies start sooner with "no clear decline in the quality of the reply" (Search Engine Journal). That's a single reported test from Anthropic, not an independent benchmark — treat it as a hypothesis to confirm against your own chat logs rather than a guarantee.
Mark every think-carefully line as a candidate for removal, but don't delete them yet. You'll test with and without in Step 6.
Step 3: Check max_tokens Caps Against Hidden Thinking
This step is the most likely to cause a production incident if skipped. If any of your Opus 5 prompts ran with thinking off and relied on a tight max_tokens cap to control cost, that cap is now a truncation risk. Anthropic's guidance notes that thinking uses part of the token budget even when you don't see the thinking output, so a cap sized for "visible output only" on Opus 5 can cut off Opus 5.5 responses mid-answer (Search Engine Journal).
Pull every max_tokens value from your configs and note which requests previously ran with thinking disabled. Those are your highest-risk candidates. You can't disable thinking on Opus 5.5 the way you could on Opus 5 at high effort or below, so any prompt built around that assumption needs its cap revisited, not just its effort level. For more on this, see read about slack code channels vs terminal agents: a decision guide.
Step 4: Review Time Budgets For Agent Workflows
If you run multiagent or agentic workflows, check whether you're passing any time-based guidance to the model. Anthropic recommends setting a time budget based on realistic task duration, and Opus 5.5 can track elapsed time against that budget internally (Search Engine Journal).
Anthropic reports that small agent groups using time signals finished research tasks faster than a single agent working without them, while holding answer quality roughly comparable to solo agents. The guide is clear that the time budget is advisory, and a hard timeout still matters for predictability, since the model may work less thoroughly when time-pressured. If your current agent setup has no time signal at all, this addition is low-risk and worth testing, since it's new capability rather than a change to existing behavior.
While you're in there, check for any tags or delimiters wrapping pasted external text like emails or scraped content. Anthropic's guidance recommends tagging pasted text with a random ID and adding a system note on how to handle it, as one layer of defense against prompt injection, not a complete one (Search Engine Journal).
Step 5: Worked Example, Before And After
Here's a simplified but representative Opus 5 system prompt for a customer support chat app, followed by the Opus 5.5 adjustment.
Before (Opus 5):
"You are a support assistant. This task involves multistep reasoning. Think carefully before responding. Effort: high. max_tokens: 800."
After (Opus 5.5, illustrative):
"You are a support assistant. Effort: medium. max_tokens: 1200."
The think-carefully line disappears because Opus 5.5 manages its own reasoning depth, with effort as the control instead. Effort drops from high to medium as the new default rather than carrying over the old high setting unexamined, matching Anthropic's reported finding that Opus 5.5 at medium effort can match or exceed Opus 5 at high effort on coding and knowledge work (Search Engine Journal). The max_tokens cap increases to account for hidden thinking tokens consuming part of the old budget, reducing truncation risk.
This is an illustrative example, not a universal template. Your actual numbers depend on your task complexity and should come out of the testing below, not be copied from here.
Step 6: Verify With Side-By-Side Testing
Before this goes live, run a fixed set of representative prompts — ideally 20 to 50 real examples pulled from production logs — through both configurations: your adjusted Opus 5.5 setup and your original Opus 5 setup. Compare three things for each pair of responses.
First, check for truncation. Did any Opus 5.5 response hit the max_tokens ceiling and cut off mid-sentence or mid-list? If so, raise the cap and retest rather than shipping a hard cutoff.
Second, compare latency. Track time to first token and total response time separately, since effort level affects both differently.
Third, run a manual quality pass on a sample of responses. Automated similarity scoring won't catch things like a shortened explanation losing important nuance.
Then repeat the effort comparison specifically: run the same prompts at medium, high, and if relevant xhigh, and decide whether the jump in effort earns back enough quality to justify the added latency and cost. Anthropic's guidance suggests saving the higher effort tiers for tasks where they demonstrably matter, rather than defaulting to them out of caution (Search Engine Journal).
Keep in mind that Anthropic's reported comparisons come from its own testing, on its own task sets. Your production traffic, prompt structure, and quality bar may not match those conditions, so the side-by-side test above is how you find out whether the guidance holds for your specific application, rather than taking the vendor's word for it.
What To Expect Next
Once you've run this checklist against your live prompts, you'll have a short list of concrete changes: likely a lower default effort setting, one or two removed think-carefully lines, an adjusted max_tokens cap, and possibly a new time budget for agent tasks. Roll these out behind a flag or to a small percentage of traffic first, and keep watching truncation rates and latency for at least a few days after the full switch. Edge cases in real traffic often surface issues that a fixed test set misses.



