- models
- Anthropic
- product
Claude Opus 5.5 in BrainBox: benchmarks, pricing and when to use it
Anthropic released Claude Opus 5.5 on September 22, 2026 and it is already selectable in the BrainBox Assistant: the top spot on the Artificial Analysis Intelligence Index, with benchmarks, intelligence units per task and when it is worth it.
Claude Opus 5.5 has been selectable in the BrainBox Assistant since September 22, 2026, the day Anthropic released it. It appears in the Pro category of the picker with a one million token context window, 128,000 max output tokens and image input. It succeeds Claude Opus 5, which stays available, and in BrainBox uses fewer intelligence units on the same task.
What is Claude Opus 5.5 and what improves over Opus 5?
It is Anthropic's frontier model for long-running agents, coding and knowledge work. Anthropic's documentation recommends starting with Opus 5.5 for most workloads and keeping Fable 5.1 for demanding reasoning. It ships with adaptive thinking always on, a one million token context window and 128,000 max output tokens.
Anthropic states that it generates output "more than 30% faster than Opus 5" and uses fewer tokens and fewer calls to finish the same work. In the benchmarks it publishes, the gain over Opus 5 is large on agent tasks:
| Benchmark, as reported by Anthropic | Opus 5 | Opus 5.5 | Claude Fable 5.1 | GPT-6 Astra |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 52.3% | 66.4% | 55.8% | 57.9% |
| FrontierCode v1.1 | 48.0% | 54.4% | 50.3% | 53.3% |
| CursorBench 4.0 | 46.6% | 57.8% | 51.8% | |
| AutomationBench | 26.9% | 40.0% | 31.4% | 41.4% |
| Humanity's Last Exam | 63.6% | 67.7% | 65.6% | 57.2% |
| OSWorld 2.0 | 74.0% | 81.8% | 80.7% | |
| Terminal-Bench-Science 0.1 | 29.0% | 58.7% | 52.6% | 64.6% |
The best score in each row is in bold.
The competitor numbers are reported by Anthropic, not by each lab, and the table itself flags standard errors of 1.6 to 5 points, so small gaps do not mean much. Anthropic also acknowledges that GPT-6 Astra stays ahead on AutomationBench and on Terminal-Bench-Science.
Independent measurement points the same way. Artificial Analysis gives it 58 on its Intelligence Index at max effort, "the highest score we have measured by several points", five above GPT-6 Astra and Claude Fable 5.1, which tie at 53, and seven above Opus 5. It leads six of the ten evaluations in the index, among them Humanity's Last Exam at 61.4% and SciCode at 66.9%, and trails on three: CritPt, AA-LCR and GDP.pdf.
How does the BrainBox Assistant work with Opus 5.5?
The BrainBox Assistant is an agent. You choose which Boxes and files it may use and which model does the reasoning; the agent decides the steps. Each one shows in the chat: "Listing sources", "Searching documents", "Reading document" with the pages it opened, "Saving note" when it sets something aside for later, and at the end the answer with page citations.
Opus 5.5 is trained for exactly that pattern. Anthropic presents it for "long and sprawling jobs like codebase-wide migrations and audits" and for financial analysis and knowledge work. In a case file review that shows up as chaining more steps without losing track of what it already read.
Open the Assistant and choose the sources
In the sidebar you tick the Boxes and files the agent may use. You can also point at a file with @ inside the message.
Click the model picker
It sits in the bottom bar of the chat. A searchable list opens.
Type 'Opus'
Claude Opus 5.5 and the earlier versions appear, with their context and the Pro and Images badges.
Choose Claude Opus 5.5
If it is missing, your personal plan does not include the Pro category or the workspace policy does not allow it.
Ask for the task
Page citations, the tools and the agent's visible steps work the same with any model.
Select model
- Claude Opus 5.5Anthropic
Latest Opus for long-running agents, coding and knowledge work
1M contextProImages - Claude Opus 5Anthropic
Previous Opus for long-running agents and professional work
1M contextProImages
BrainBox is an AI workspace for documents that answers with citations to the exact page. The model changes how the answer is reasoned; the trail back to the source does not.
How many intelligence units does Opus 5.5 use per task?
BrainBox does not measure usage in tokens but in intelligence units, and each task uses a different number depending on how many searches and reads the agent makes, how much the model reasons and how much it writes. Estimates for the standard plan, with GPT-6 Sol alongside for reference:
| Task in the Assistant | Opus 5.5 | Opus 5 | GPT-6 Sol |
|---|---|---|---|
| Single cited question, one or two searches | ≈ 7 units | ≈ 8 units | ≈ 4 units |
| Spreadsheet analysis with the code interpreter | ≈ 22 units | ≈ 27 units | ≈ 11 units |
| Case file review with about ten searches and reads | ≈ 52 units | ≈ 65 units | ≈ 27 units |
| Full read of a 200 page contract, page by page | ≈ 57 units | ≈ 71 units | ≈ 29 units |
| PDF report built from a case file, with code and about fifteen steps | ≈ 95 units | ≈ 118 units | ≈ 49 units |
These are estimates: a longer answer, more sources or more agent steps push the number up, and Opus 5.5 tends toward long answers. For reference, the Elite plan includes 400 units a month, so one case file review a day on Opus 5.5 takes a good share of the month.
What happens if the Box is larger than the context window?
Nothing is left out. The Assistant does not load the whole Box into the model: it searches the selected sources and reads the pages that answer the question, each with its file and page. It works the same for a Box of 200 pages or 20,000. The model decides how to reason over what it reads; the size of the Box does not decide which model you can pick.
Which tasks is it worth choosing for?
Opus 5.5 is the most capable model in the picker and the heaviest. What to decide on each task is whether that difference in usage pays for itself.
| Task in the Assistant | Opus 5.5? | Alternative |
|---|---|---|
| Single cited question | Not needed; a default model answers the same with far fewer units | GPT-6 Luna |
| An opinion or review where being wrong is expensive | Yes, it is the highest score on the independent index | GPT-6 Sol if budget rules |
| Comparing contract versions with many addenda | Yes, it chains more steps without losing the thread | GPT-6 Sol |
| PDF report or data analysis with the code interpreter | Yes, if the report will be defended before a committee | GPT-6 Sol for drafts |
| Batch summaries or document sorting | No, the usage is not justified | GPT-6 Luna |
| Questions about screenshots or diagrams | Yes, it accepts images | Any other model with the Images badge |
A legal team can keep Opus 5.5 enabled only for final opinions and work day to day on a default model, from the workspace model policy. Citations still point to the page with any of them, so verifying the answer works the same.
What about safety?
Anthropic reports that in its containment evaluation, Opus 5.5 "attempted to circumvent boundaries around 85% less often than Opus 5", that it is the strongest-performing model in its alignment audit to date, and that it matches or beats Opus 5 on prompt injection resistance in every setting they tested. These are the company's own evaluations and they deliberately test hard situations, not typical use. In the Assistant the practical protection is the same as always: the agent asks for approval before sensitive actions on your files, and every claim carries its citation.
Anthropic says Opus 5.5 completed a 680,000-line code migration in less than a day; in the Assistant, a long review can take several minutes and you watch it advance step by step in the chat.
Sources
- Anthropic, Introducing Claude Opus 5.5: announcement of September 22, 2026, benchmark table, speed, efficiency and safety evaluations.
- Anthropic, models overview: one million token context, max output, adaptive thinking and usage recommendation.
- Artificial Analysis, Claude Opus 5.5 takes the top spot: Intelligence Index, evaluations it leads, output tokens per task and comparison with GPT-6 Astra and Fable 5.1.
- Artificial Analysis, Claude Opus 5.5 model page: score per effort level, context, modalities and verbosity.
- Artificial Analysis, model leaderboard: scores for GPT-6 Astra, GPT-6 Sol, Claude Fable 5.1 and the Opus 5.5 effort levels.
Frequently asked questions
- What changed between Claude Opus 5 and Opus 5.5?
- According to Anthropic, Opus 5.5 performs at roughly the level of Claude Fable 5.1 on most tasks, generates output more than 30% faster than Opus 5, and uses fewer tokens and fewer calls to finish the same work. On Terminal-Bench 4.0 it goes from 52.3% to 66.4%, and on Humanity's Last Exam from 63.6% to 67.7%. On the Artificial Analysis Intelligence Index it moves from 51 to 58. In BrainBox the same task uses about 20% fewer intelligence units.
- How much does Opus 5.5 cost in BrainBox?
- Usage is measured in intelligence units, not tokens. Estimates on the standard plan: a single cited question, about 7 units; a data analysis with the code interpreter, about 22; a case file review with searches and page reads, about 52; a PDF report built from a case file, about 95. It is the top tier in the picker, and for most questions a default model uses far less.
- Which plans include it?
- Opus 5.5 sits in the Pro category of the picker. On personal accounts it shows up on paid plans; in a workspace it depends on the model policy the admin has set. If you cannot see it, check the workspace AI policy.
- Is it better than GPT-6 for my documents?
- On the Artificial Analysis Intelligence Index, Opus 5.5 scores 58 against 53 for GPT-6 Astra and 48 for GPT-6 Sol. It leads six of the ten evaluations in the index. The tradeoff is usage: Artificial Analysis calls it very verbose, and in BrainBox a long review uses close to twice the units of GPT-6 Sol.
- What if my Box is larger than the context window?
- Nothing is left out. The Assistant does not load the whole Box at once: it searches the sources you chose and reads the pages that answer the question, each with its citation. The size of the Box does not limit which model you can use.
- Is Claude Opus 5 still available?
- Yes. Opus 5 stays in the picker and uses somewhat more intelligence units per task than Opus 5.5. Chats that were already using Opus 5 do not switch models on their own.
Written by
BrainBox Team
Document intelligence, by ExaByte Company
We build BrainBox — the platform teams use to ask questions across their own documents and get answers with exact page citations.
Keep reading
GPT-6 Sol vs Luna: benchmarks, pricing and which to use in BrainBox
OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026 and both are already in the BrainBox Assistant: benchmarks, Artificial Analysis score, intelligence units per task and which one to pick.
Grok 4.7 in BrainBox: benchmarks, pricing and when to use it
xAI's Grok 4.7 is now selectable in the BrainBox Assistant: benchmarks against Grok 4.6, its Artificial Analysis score, the intelligence units it uses per task, and which document work it suits.
Xiaomi MiMo V2.6 Pro vs Flash: benchmarks, pricing and which to use in BrainBox
Xiaomi's MiMo V2.6 Pro and Flash, MIT open weights with a 1M context window, are now in the BrainBox Assistant: benchmarks, Artificial Analysis score, intelligence units per task and which to pick.