Claude Opus 5: Everything you need to know

Written by
Last updated on:
August 7, 2026
Written by
Last updated on:
August 7, 2026

Anthropic's newest Opus-tier model closes the gap with its most capable system, Claude Fable 5, while holding the line on price—reshaping what teams can expect from everyday AI work.

Abstract illustration for Anthropic’s Claude Opus 5 announcement, with speckled pebble-like forms arranged in a vertical cluster on a warm beige background.

Claude Opus 5 is Anthropic’s newest Opus‑tier model, built for serious coding, long‑running agents, and complex enterprise knowledge work. Described by Anthropic as delivering near‑Fable‑level intelligence on many real‑world tasks while keeping Opus‑tier pricing, the model is intended to serve as a primary model for production‑grade workloads rather than a niche frontier option.

With Opus 5, Anthropic is setting a new baseline for the Claude 5 family. It succeeds Opus 4.8 in the Opus tier and is now Anthropic’s recommended choice for demanding, day‑to‑day use in Claude Max and Claude Pro. If your organization already uses Claude for development, automation, or research, Opus 5 is the version Anthropic recommends for ongoing, enterprise‑scale use—a more capable and reliable successor that remains cost‑effective in practice.

Performance and cost: frontier capability at Opus prices

Anthropic positions Opus 5 as a major performance jump over Opus 4.8 at the same price point: $5 per million input tokens and $25 per million output tokens, with much stronger results across software engineering, knowledge work, and problem‑solving benchmarks.

On coding and knowledge‑work evaluations, it sets new state‑of‑the‑art scores while staying behind Mythos‑class models only on the riskiest cybersecurity tasks.

  • On Frontier‑Bench v0.1, Opus 5 more than doubles Opus 4.8’s performance at a lower cost per task, and surpasses all other models tested.
  • On CursorBench 3.2, at max effort, Opus 5 performs within 0.5% of Fable 5’s peak score—but at half the cost per task. At high, xhigh, and max effort, it delivers better performance at a given cost than any competing model.
  • On ARC‑AGI 3, which tests how models solve novel problems, Opus 5’s score is three times as high as the next‑best model.
  • On Zapier AutomationBench, measuring end‑to‑end business task completion, Opus 5’s pass rate is around 1.5× the next‑best model at the same cost per task, and even at its lowest effort setting it passes more tasks than any other model.
  • On OSWorld 2.0, a computer‑use benchmark, Opus 5 outperforms every other model at any given cost, surpassing Fable 5’s best result at just over a third of the cost.

Across evaluations like CursorBench and Zapier AutomationBench, Opus 5 maintains strong results even at lower effort settings and scales smoothly as you increase effort, giving teams a practical way to tune cost and latency.

Line chart titled “Agentic coding by effort level” comparing Frontier‑Bench v0.1 scores against cost per attempt for Opus 5, Fable 5, Opus 4.8, and GPT‑5.6 Sol. Opus 5 reaches the highest score, around 44%, at roughly $15 per attempt.

What Opus 5 can do

Stronger agency and long-horizon work

Opus 5 stands out most clearly on long‑horizon tasks, where it verifies its own work and iterates until it reaches a solution. 

  • Rebuilding a 3D part from raw pixels. On a Frontier‑Bench task, Opus 5 was given only a drawing of a machine part and asked to rebuild it as a 3D FreeCAD model — without any way to view the drawing directly. Opus 5 wrote its own computer vision pipeline to extract geometry from the pixels and repeatedly reconstructed the part. No competing model could solve the task after five attempts.
  • Fixing a real-world bug more thoroughly than the community patch. Given a real bug in a popular open‑source package manager, Opus 5 found the root cause and fixed an edge case that the community patch had missed. Another model fixed only the surface symptom and mistakenly reported success.
  • Building a market data feed in one session. An engineer at a trading firm used Opus 5 to build a new market data feed end‑to‑end. When no live feed was available to test against, Opus 5 built its own test harness to validate its parsing of the exchange’s data correctly — a task previous models couldn’t complete even with extensive plans.

For engineering teams, this translates directly into more dependable long‑horizon behavior: Opus 5 is more likely to keep working through obstacles, verify its own assumptions, and reach a usable solution without extensive hand‑holding.

Production-grade coding and debugging

Opus 5 is positioned first and foremost as Anthropic’s go‑to model for serious coding work. Rather than only generating snippets, it’s designed to support production use cases like automated code review, feature development, and bug‑fixing, with outputs that teams can realistically integrate into their codebases.

“Claude Opus 5 came out ahead of every model in its family on our internal evals,” says Fabian Hedin, co-founder of Lovable. “It isn’t just better on our hardest coding tasks, up 22% over Opus 4.7, it’s steadier, with far less variance run to run.”

Together, these traits make Opus 5 a stronger fit for production‑grade tasks such as automated PR review and code review assistance, long‑running refactors and migration work, and CI helpers that diagnose failing tests and propose fixes.

Read our benchmarks here if you're interested in seeing how Opus 4.7 matches against Sonnet 5.

Better knowledge work, research, and analysis

Opus 5 is also built for complex knowledge work, such as document‑heavy enterprise tasks and scientific and financial analysis. Anthropic describes it as a “meaningful improvement” over Opus 4.8 for scientific research, with better performance across life sciences tasks such as structural biology, organic chemistry, and bioinformatics.

“On our genomics analysis work, Claude Opus 5 behaves more like a careful scientist than any model we’ve run,” says Alfredo Andere, CEO of LatchBio. “It reaches for the right statistical tests to rule out confounders, cross‑checks its own results by independent methods, and stays on track through long multi‑step analyses.”

Claude Opus 5’s strengths apply to other forms of knowledge work, too. In document-heavy enterprise workflows, Opus 5 can pull together information across long files, surface the details that matter, and organize its findings into something people can review and use. In financial work, that can mean following the logic of a model or forecast without losing the context around the numbers.

Visual understanding and artifact quality

Opus 5 handles visual inputs more reliably than earlier models, reading charts, documents, and diagrams and turning that understanding into well‑structured slides or written reports. For teams shipping dashboards, marketing sites, or internal decks, this combination of stronger visual understanding and higher‑quality output makes Opus 5 a more capable partner whenever layout and presentation matter.

“Claude Opus 5 checks its own work the way a real frontend developer would,” says AJ Orbach, co-founder and CEO of Triple Whale, an e-commerce intelligence and analytics platform. “On our benchmark it opened its pages in a browser at desktop and phone widths, caught a product hidden below the mobile fold and an off‑screen checkout button, and fixed both before handing the work back.”

Benchmark table comparing Opus 5, Fable 5, Opus 4.8, and GPT‑5.6 Sol across coding, knowledge work, reasoning, search, computer use, business, legal, health, and biology tasks. Opus 5 is highlighted and leads several benchmarks, including terminal coding, knowledge work, novel problem-solving, computer use, and business workflows.

Alignment, safety, and guardrails

Anthropic makes it clear that Opus 5 does not push the frontier on risky, dual‑use capabilities. Instead, they frame it as their most aligned and safest generally available model so far.

Alignment and behavior

Opus 5 is described as Anthropic’s most aligned model so far. In an automated behavioral audit that tests for issues like cooperation with misuse and deceptive behavior, it scored 2.3 on overall misaligned behavior—the lowest of Anthropic’s recent models and below Opus 4.8, Sonnet 5, and Fable 5. It shows the strongest adherence to Claude’s Constitution, the lowest rates of deceptive responses, and is the least susceptible to being tricked into misuse among the models evaluated.

Anthropic also calls Opus 5 its safest model yet when it comes to avoiding reckless actions with hard‑to‑reverse consequences. In practice, that means more cautious judgment and better self‑checking: the model is less likely to push ahead with risky instructions, more likely to flag uncertainty or ask for clarification, and generally more conservative when the stakes are high. These changes make it useful for teams deploying long‑running agents with access to tools, code, or sensitive data.

Cybersecurity and biology safeguards

Opus 5’s safeguards are meant to enable useful cybersecurity and biology work while keeping tighter controls on high‑risk behavior. In security contexts, it can help teams find vulnerabilities in source code for secure development and code review, while classifiers continue to block binary‑based scanning, penetration testing, and exploit generation. Those classifiers are expected to intervene far less often than Fable 5’s, which makes Opus 5 more practical for everyday use while still enforcing limits where needed.

Effort, Fast mode, and tooling updates

Opus 5 works within Anthropic’s newer effort and Fast mode paradigm, giving teams more control over speed, cost, and capability.

Effort and test-time compute

Anthropic’s broader Claude 5 documentation shows that Opus 5 supports the full effort ladder—low, medium, high, xhigh, and max—with thinking on by default and max as the top tier. Customers can tune effort to:

  • Scale intelligence up for harder tasks (e.g., max effort on complex debugging or research).
  • Conserve tokens and latency for routine work (e.g., low or medium effort on straightforward summarization or CRUD code).

In the announcement, Anthropic explicitly plots performance vs. effort and cost to show how Opus 5 can be tuned for either maximum capability or cost‑efficient throughput.

Fast mode and pricing

Opus 5 is offered in Fast mode, running around 2.5× the default speed.

  • Base pricing: $5 per million input tokens, $25 per million output tokens—same as Opus 4.8.
  • Fast mode: Available at twice Opus 5’s base price on the Claude Platform and through usage credits in Claude Code.

This lets teams run latency‑sensitive workloads (e.g., interactive IDE copilots, live chat support) without giving up the benefits of Opus‑class reasoning.

Developer experience updates

Alongside Opus 5, Anthropic is shipping two platform updates in beta:

  • Mid‑conversation tool changes on Claude Platform. You can now change which tools Claude can use mid‑conversation without invalidating the prompt cache. This matters for long‑running agents that need to adapt their toolset as they work.
  • Automatic fallbacks on the API. Customers can configure requests flagged by safety classifiers on Opus 5 (or Fable 5) to automatically route to another model. With fallbacks enabled, API calls default to “best available model” instead of being blocked outright.

Opus 5 also continues the Opus line’s data‑handling stance: it has no data retention requirements for general access, consistent with earlier models.

Getting started with Claude Opus 5

Claude Opus 5 is available across Anthropic’s ecosystem at the same price as Opus 4.8, including in Claude Max and Claude Pro, via the Claude API, and through cloud partners like Amazon Bedrock and Google Cloud.

If you already use Claude for coding or knowledge work, a practical first step is to move existing Opus 4.8 workflows over to Opus 5 and compare accuracy, latency, and cost. From there, you can start tuning effort levels for different task types and pilot agents that take advantage of Opus 5’s stronger reasoning and analytical capabilities.

For teams still exploring AI use cases, Opus 5 offers a sensible default for development tools, research assistants, and document‑heavy processes—with Mythos‑class models like Fable 5 reserved for situations where you truly need the highest frontier performance.

If you’re interested in seeing how Opus 4.8, an earlier model, compares to OpenAI’s Terra, Sol, and Luna models, you can read our benchmarks here.

Learn more

Frequently Asked Questions

Claude Opus 5 is Anthropic's newest Opus-tier model, built for serious coding, long-running agents, and complex enterprise knowledge work. It succeeds Opus 4.8 and delivers near-Fable-level intelligence on many real-world tasks while keeping the same Opus-tier pricing.

Claude Opus 5 is priced at $5 per million input tokens and $25 per million output tokens, the same rate as Opus 4.8. A Fast mode option is also available at twice that base price for latency-sensitive workloads.

Opus 5 performs close to Fable 5 on many benchmarks, coming within 0.5% of its peak CursorBench 3.2 score at half the cost per task, though it remains behind Mythos-class models on the riskiest cybersecurity tasks. Fable 5 is still the better choice for situations that demand the highest frontier performance.

Yes. Opus 5's safeguards allow it to help find vulnerabilities in source code for secure development and code review, while classifiers block higher-risk activities like binary-based scanning, penetration testing, and exploit generation. These classifiers intervene far less often than Fable 5's, making Opus 5 more practical for everyday security-adjacent work.

Claude Opus 5 is available through Claude Max, Claude Pro, the Claude API, and cloud partners including Amazon Bedrock and Google Cloud, all at the same price as Opus 4.8.