Skip to main content
  • models
  • Anthropic
  • product

Claude Opus 5.5 in BrainBox: benchmarks, pricing and when to use it

Anthropic released Claude Opus 5.5 on September 22, 2026 and it is already selectable in the BrainBox Assistant: the top spot on the Artificial Analysis Intelligence Index, with benchmarks, intelligence units per task and when it is worth it.

BrainBox Team8 min read
View as Markdown

Claude Opus 5.5 has been selectable in the BrainBox Assistant since September 22, 2026, the day Anthropic released it. It appears in the Pro category of the picker with a one million token context window, 128,000 max output tokens and image input. It succeeds Claude Opus 5, which stays available, and in BrainBox uses fewer intelligence units on the same task.

What is Claude Opus 5.5 and what improves over Opus 5?

It is Anthropic's frontier model for long-running agents, coding and knowledge work. Anthropic's documentation recommends starting with Opus 5.5 for most workloads and keeping Fable 5.1 for demanding reasoning. It ships with adaptive thinking always on, a one million token context window and 128,000 max output tokens.

Anthropic states that it generates output "more than 30% faster than Opus 5" and uses fewer tokens and fewer calls to finish the same work. In the benchmarks it publishes, the gain over Opus 5 is large on agent tasks:

Benchmark, as reported by AnthropicOpus 5Opus 5.5Claude Fable 5.1GPT-6 Astra
Terminal-Bench 4.052.3%66.4%55.8%57.9%
FrontierCode v1.148.0%54.4%50.3%53.3%
CursorBench 4.046.6%57.8%51.8%
AutomationBench26.9%40.0%31.4%41.4%
Humanity's Last Exam63.6%67.7%65.6%57.2%
OSWorld 2.074.0%81.8%80.7%
Terminal-Bench-Science 0.129.0%58.7%52.6%64.6%

The best score in each row is in bold.

The competitor numbers are reported by Anthropic, not by each lab, and the table itself flags standard errors of 1.6 to 5 points, so small gaps do not mean much. Anthropic also acknowledges that GPT-6 Astra stays ahead on AutomationBench and on Terminal-Bench-Science.

Independent measurement points the same way. Artificial Analysis gives it 58 on its Intelligence Index at max effort, "the highest score we have measured by several points", five above GPT-6 Astra and Claude Fable 5.1, which tie at 53, and seven above Opus 5. It leads six of the ten evaluations in the index, among them Humanity's Last Exam at 61.4% and SciCode at 66.9%, and trails on three: CritPt, AA-LCR and GDP.pdf.

How does the BrainBox Assistant work with Opus 5.5?

The BrainBox Assistant is an agent. You choose which Boxes and files it may use and which model does the reasoning; the agent decides the steps. Each one shows in the chat: "Listing sources", "Searching documents", "Reading document" with the pages it opened, "Saving note" when it sets something aside for later, and at the end the answer with page citations.

Opus 5.5 is trained for exactly that pattern. Anthropic presents it for "long and sprawling jobs like codebase-wide migrations and audits" and for financial analysis and knowledge work. In a case file review that shows up as chaining more steps without losing track of what it already read.

  1. Open the Assistant and choose the sources

    In the sidebar you tick the Boxes and files the agent may use. You can also point at a file with @ inside the message.

  2. Click the model picker

    It sits in the bottom bar of the chat. A searchable list opens.

  3. Type 'Opus'

    Claude Opus 5.5 and the earlier versions appear, with their context and the Pro and Images badges.

  4. Choose Claude Opus 5.5

    If it is missing, your personal plan does not include the Pro category or the workspace policy does not allow it.

  5. Ask for the task

    Page citations, the tools and the agent's visible steps work the same with any model.

Select model

Select model

Opus
  • Anthropic logo
    Claude Opus 5.5Anthropic

    Latest Opus for long-running agents, coding and knowledge work

    1M contextProImages
  • Anthropic logo
    Claude Opus 5Anthropic

    Previous Opus for long-running agents and professional work

    1M contextProImages
Simulation of the BrainBox Assistant model picker with Claude Opus 5.5 selected.

BrainBox is an AI workspace for documents that answers with citations to the exact page. The model changes how the answer is reasoned; the trail back to the source does not.

How many intelligence units does Opus 5.5 use per task?

BrainBox does not measure usage in tokens but in intelligence units, and each task uses a different number depending on how many searches and reads the agent makes, how much the model reasons and how much it writes. Estimates for the standard plan, with GPT-6 Sol alongside for reference:

Task in the AssistantOpus 5.5Opus 5GPT-6 Sol
Single cited question, one or two searches≈ 7 units≈ 8 units≈ 4 units
Spreadsheet analysis with the code interpreter≈ 22 units≈ 27 units≈ 11 units
Case file review with about ten searches and reads≈ 52 units≈ 65 units≈ 27 units
Full read of a 200 page contract, page by page≈ 57 units≈ 71 units≈ 29 units
PDF report built from a case file, with code and about fifteen steps≈ 95 units≈ 118 units≈ 49 units

These are estimates: a longer answer, more sources or more agent steps push the number up, and Opus 5.5 tends toward long answers. For reference, the Elite plan includes 400 units a month, so one case file review a day on Opus 5.5 takes a good share of the month.

What happens if the Box is larger than the context window?

Nothing is left out. The Assistant does not load the whole Box into the model: it searches the selected sources and reads the pages that answer the question, each with its file and page. It works the same for a Box of 200 pages or 20,000. The model decides how to reason over what it reads; the size of the Box does not decide which model you can pick.

Which tasks is it worth choosing for?

Opus 5.5 is the most capable model in the picker and the heaviest. What to decide on each task is whether that difference in usage pays for itself.

Task in the AssistantOpus 5.5?Alternative
Single cited questionNot needed; a default model answers the same with far fewer unitsGPT-6 Luna
An opinion or review where being wrong is expensiveYes, it is the highest score on the independent indexGPT-6 Sol if budget rules
Comparing contract versions with many addendaYes, it chains more steps without losing the threadGPT-6 Sol
PDF report or data analysis with the code interpreterYes, if the report will be defended before a committeeGPT-6 Sol for drafts
Batch summaries or document sortingNo, the usage is not justifiedGPT-6 Luna
Questions about screenshots or diagramsYes, it accepts imagesAny other model with the Images badge

A legal team can keep Opus 5.5 enabled only for final opinions and work day to day on a default model, from the workspace model policy. Citations still point to the page with any of them, so verifying the answer works the same.

What about safety?

Anthropic reports that in its containment evaluation, Opus 5.5 "attempted to circumvent boundaries around 85% less often than Opus 5", that it is the strongest-performing model in its alignment audit to date, and that it matches or beats Opus 5 on prompt injection resistance in every setting they tested. These are the company's own evaluations and they deliberately test hard situations, not typical use. In the Assistant the practical protection is the same as always: the agent asks for approval before sensitive actions on your files, and every claim carries its citation.

Anthropic says Opus 5.5 completed a 680,000-line code migration in less than a day; in the Assistant, a long review can take several minutes and you watch it advance step by step in the chat.

Sources

  • Anthropic, Introducing Claude Opus 5.5: announcement of September 22, 2026, benchmark table, speed, efficiency and safety evaluations.
  • Anthropic, models overview: one million token context, max output, adaptive thinking and usage recommendation.
  • Artificial Analysis, Claude Opus 5.5 takes the top spot: Intelligence Index, evaluations it leads, output tokens per task and comparison with GPT-6 Astra and Fable 5.1.
  • Artificial Analysis, Claude Opus 5.5 model page: score per effort level, context, modalities and verbosity.
  • Artificial Analysis, model leaderboard: scores for GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1 and the Opus 5.5 effort levels.

Frequently asked questions

What changed between Claude Opus 5 and Opus 5.5?
According to Anthropic, Opus 5.5 performs at roughly the level of Claude Fable 5.1 on most tasks, generates output more than 30% faster than Opus 5, and uses fewer tokens and fewer calls to finish the same work. On Terminal-Bench 4.0 it goes from 52.3% to 66.4%, and on Humanity's Last Exam from 63.6% to 67.7%. On the Artificial Analysis Intelligence Index it moves from 51 to 58. In BrainBox the same task uses about 20% fewer intelligence units.
How much does Opus 5.5 cost in BrainBox?
Usage is measured in intelligence units, not tokens. Estimates on the standard plan: a single cited question, about 7 units; a data analysis with the code interpreter, about 22; a case file review with searches and page reads, about 52; a PDF report built from a case file, about 95. It is the top tier in the picker, and for most questions a default model uses far less.
Which plans include it?
Opus 5.5 sits in the Pro category of the picker. On personal accounts it shows up on paid plans; in a workspace it depends on the model policy the admin has set. If you cannot see it, check the workspace AI policy.
Is it better than GPT-6 for my documents?
On the Artificial Analysis Intelligence Index, Opus 5.5 scores 58 against 53 for GPT-6 Astra and 48 for GPT-6 Sol. It leads six of the ten evaluations in the index. The tradeoff is usage: Artificial Analysis calls it very verbose, and in BrainBox a long review uses close to twice the units of GPT-6 Sol.
What if my Box is larger than the context window?
Nothing is left out. The Assistant does not load the whole Box at once: it searches the sources you chose and reads the pages that answer the question, each with its citation. The size of the Box does not limit which model you can use.
Is Claude Opus 5 still available?
Yes. Opus 5 stays in the picker and uses somewhat more intelligence units per task than Opus 5.5. Chats that were already using Opus 5 do not switch models on their own.

Written by

BrainBox Team

Document intelligence, by ExaByte Company

We build BrainBox — the platform teams use to ask questions across their own documents and get answers with exact page citations.