- models
- xAI
- product
Grok 4.7 in BrainBox: benchmarks, pricing and when to use it
xAI's Grok 4.7 is now selectable in the BrainBox Assistant: benchmarks against Grok 4.6, its Artificial Analysis score, the intelligence units it uses per task, and which document work it suits.
Grok 4.7 has been selectable in the BrainBox Assistant since September 21, 2026, the day xAI released it. It appears in the Pro category of the model picker with a 500K token window and image input. It succeeds Grok 4.6, which stays available, and in BrainBox uses fewer intelligence units on the same task.
What is Grok 4.7 and what improves over Grok 4.6?
Grok 4.7 is xAI's flagship model for coding, agents and knowledge work. According to xAI's announcement, it "uses a new, larger base model compared to Grok 4.6", was "trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete", and is "better at verifying its own work". The context window stays at 500K tokens, unchanged from Grok 4.6.
In the benchmarks xAI publishes, Grok 4.7 improves on every test over Grok 4.6, beats Claude Fable 5.1 on EEBench and on the Harvey legal benchmark, and trails it on CursorBench, Terminal-Bench and HealthBench. The competitor numbers are also xAI's, not each lab's.
| Benchmark, as reported by xAI | Grok 4.6 | Grok 4.7 | Claude Fable 5.1 | GPT-5.6 Sol |
|---|---|---|---|---|
| DeepSWE v1.1, software agent | 65.2% | 71.0% | 70.0% | 72.7% |
| CursorBench 4.0, coding in the editor | 40.4% | 46.3% | 51.8% | 41.7% |
| Terminal-Bench 4.0, terminal agent | 20.3% | 38.0% | 57.9% | 37.3% |
| HealthBench Professional | 48.5% | 56.7% | 62.1% | 60.5% |
| EEBench, electrical engineering | 53.0% | 64.0% | 56.4% | 39.4% |
| Harvey Legal Agent Benchmark | 15.8% | 19.6% | 6.7% | 2.5% |
Artificial Analysis, which measures models independently, gives it 46 on its Intelligence Index v4.3, two points above Grok 4.6. That puts xAI among the four highest scoring labs, but seven points behind Claude Fable 5.1 and GPT-6 Astra, which score 53 in their maximum setting. On the same report's Coding Agent Index, Grok 4.7 with Grok Build scores 56, fourth behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5.
Grok 4.7 is talkative: according to the same report, it generates about 81,000 output tokens per index task, against 36,000 for Grok 4.6, so it reasons more and one of its answers uses more units. It also hallucinates less: on AA-Omniscience, the Artificial Analysis knowledge test, its hallucination rate drops from 34% to 29%, and accuracy almost unchanged, 47% against 48%.
How does the BrainBox Assistant work with Grok 4.7?
The BrainBox Assistant is an agent. You choose which Boxes and files it may use and which model does the reasoning; the agent decides the steps. Each one shows in the chat: "Listing sources", "Searching documents", "Reading document" with the pages it opened, "Saving note" when it sets something aside for later, and at the end the answer with page citations. For a large case file it does not load everything at once: it searches, reads the pages that matter, and if the task calls for it, reads a whole document.
Grok 4.7 was trained for tasks that chain many steps, and that is what the Assistant does in a review: several searches, several reads, a cross-check between documents and a long answer. That is where the gain over Grok 4.6 shows.
How do you pick Grok 4.7 in BrainBox?
Open the Assistant and choose the sources
In the sidebar you tick the Boxes and files the agent may use. You can also point at a file with @ inside the message.
Click the model picker
It sits in the bottom bar of the chat. A searchable list opens.
Type 'Grok'
Grok 4.7 and Grok 4.6 appear with their context and the Pro and Images badges.
Choose Grok 4.7
If it is missing, your personal plan does not include the Pro category or the workspace policy does not allow it.
Ask for the task
Page citations, the tools and the agent's visible steps work the same with any model.
Select model
- Grok 4.7xAI
Long-running coding, agents and knowledge work
500K contextProImages - Grok 4.6xAI
Coding, knowledge work and STEM; succeeded by Grok 4.7
500K contextProImages
BrainBox is an AI workspace for documents that answers with citations to the exact page. The model changes how the answer is reasoned; the trail back to the source does not.
How many intelligence units does Grok 4.7 use per task?
BrainBox does not measure usage in tokens but in intelligence units, and each task uses a different number depending on how many searches and reads the agent makes, how much the model reasons and how much it writes. Estimates for the standard plan:
| Task in the Assistant | Grok 4.7 | Grok 4.6 |
|---|---|---|
| Single cited question, one or two searches | ≈ 3 units | ≈ 3 units |
| Spreadsheet analysis with the code interpreter | ≈ 8 units | ≈ 9 units |
| Case file review with about ten searches and reads | ≈ 20 units | ≈ 25 units |
| Full read of a 200 page contract, page by page | ≈ 22 units | ≈ 28 units |
| PDF report built from a case file, with code and about fifteen steps | ≈ 35 units | ≈ 42 units |
| A 450 page case file read whole in a single pass | ≈ 95 units | ≈ 120 units |
These are estimates: a longer answer, more selected sources or more agent steps push the number up. For reference, the Elite plan includes 400 units a month.
What happens if the Box is larger than Grok 4.7's context?
Nothing is left out. The Assistant does not load the whole Box into the model: it searches the selected sources and reads the pages that answer the question, each with its file and page. It works the same for a Box of 200 pages or 20,000. The model decides how to reason over what it reads; the size of the Box does not decide which model you can pick.
Which tasks is it worth choosing for?
Grok 4.7 fits multi-step tasks: reviewing a case file of dozens of documents, comparing two versions of a contract, cross-checking a tender against its addenda, or asking for a report delivered as a PDF. On the Harvey Legal Agent Benchmark, as reported by xAI, Grok 4.7 scores 19.6% against 15.8% for Grok 4.6, 6.7% for Claude Fable 5.1 and 2.5% for GPT-5.6 Sol. That is a low number in absolute terms for every model. A legal memo drafted with Grok 4.7 gets reviewed the same way as one drafted with any other model.
| Task in the Assistant | Grok 4.7? | Alternative |
|---|---|---|
| Single cited question | Not needed; a default model answers the same with fewer units | Default or Flash |
| Full case file review | Yes, long multi-step work is its strength | Claude or GPT if budget allows |
| Data analysis with the code interpreter | Yes, better at coding than 4.6 | Claude Fable 5.1 leads on coding agents |
| Questions about screenshots or diagrams | Yes, it accepts images | Any other model with the Images badge |
| Reading more than 300 pages in a single pass | Carefully, unit usage nearly doubles | Xiaomi MiMo V2.6, 1M context with no usage cliff |
You can test it inside the same conversation: switch from Grok 4.6 to 4.7 in the picker, repeat the question and compare. Citations still point to the page, so verifying the answer works the same with both.
How is this different from using Grok in xAI's app?
The Grok app, Cursor and Grok Build give access to the same model, according to xAI. What differs is everything around it.
| Grok app | BrainBox with Grok 4.7 | |
|---|---|---|
| Sources | Whatever you paste or attach in that conversation | The Box's indexed documents: PDF, Office, transcribed audio, images |
| Citations | Do not point to a page in your files | Page, passage and expanded context for every claim |
| Switching models | xAI models only | Grok 4.7 next to Claude, GPT, Gemini and open models, in the same chat |
| Team | Individual account | Shared Box with roles; the admin decides which models are allowed |
| Usage | Flat subscription | Intelligence units per use, based on the model you pick |
For legal teams, the workspace model policy lets you enable Grok 4.7 for the team or block it, without touching each account. The rest of the document workspace works the same with any model.
A full task on the Artificial Analysis index took Grok 4.7 about 7.1 minutes on average; in BrainBox, a long review can take several minutes and you watch it advance step by step in the chat.
Sources
- xAI, Introducing Grok 4.7: announcement of September 21, 2026, benchmark table and availability.
- Artificial Analysis, Benchmarking Grok 4.7: Intelligence Index, Coding Agent Index, tokens per task, AA-Omniscience, 500K context.
- Artificial Analysis, model leaderboard: scores for Claude Fable 5.1, GPT-6 Astra and Grok 4.6.
- Artificial Analysis, Grok 4.7 model page: context and modalities.
Frequently asked questions
- What changed between Grok 4.6 and Grok 4.7?
- According to xAI, Grok 4.7 uses a larger base model, a longer reinforcement learning run on tasks that take hours, and checks its own output more carefully. On DeepSWE v1.1 it goes from 65.2% to 71.0%; on the Harvey Legal Agent Benchmark, from 15.8% to 19.6%. On the Artificial Analysis Intelligence Index it moves from 44 to 46. The context window stays at 500K tokens. In BrainBox the same task also uses 10% to 20% fewer intelligence units than with Grok 4.6.
- How much does Grok 4.7 cost in BrainBox?
- Usage is measured in intelligence units, not tokens. Estimates on the standard plan: a single cited question, about 3 units; a data analysis with the code interpreter, about 8; a case file review with searches and page reads, about 20; a PDF report built from a case file, about 35. Each uses 10% to 20% fewer units than Grok 4.6.
- Which plans include it?
- Grok 4.7 sits in the Pro category of the picker. On personal accounts it shows up on paid plans; in a workspace it depends on the model policy the admin has set. If you cannot see it, check the workspace AI policy.
- Does Grok 4.7 read images?
- Yes. In BrainBox it is flagged with image support, so you can attach a screenshot or a diagram to a question, or ask about the images indexed in the Box.
- Is it better than Claude or GPT for my documents?
- On the Artificial Analysis Intelligence Index, Claude Fable 5.1 and GPT-6 Astra score 53 against 46 for Grok 4.7. It also trails on coding agents. What it has going for it is usage: a long case file review uses noticeably fewer intelligence units than with those models. On a single cited question, neither the quality gap nor the usage gap shows.
- What if my Box is larger than Grok 4.7's context window?
- Nothing is left out. The Assistant does not load the whole Box at once: it searches the sources you chose and reads the pages that answer the question, each with its citation. The size of the Box does not limit which model you can use.
- Is Grok 4.6 still available?
- Yes. Grok 4.6 stays in the picker and uses somewhat more intelligence units per task. Chats that were already using 4.6 do not switch models on their own.
Written by
BrainBox Team
Document intelligence, by ExaByte Company
We build BrainBox — the platform teams use to ask questions across their own documents and get answers with exact page citations.
Keep reading
Xiaomi MiMo V2.6 Pro vs Flash: benchmarks, pricing and which to use in BrainBox
Xiaomi's MiMo V2.6 Pro and Flash, MIT open weights with a 1M context window, are now in the BrainBox Assistant: benchmarks, Artificial Analysis score, intelligence units per task and which to pick.
BrainBox in 2026: the workspace for professionals who work with documents
Research tools keep you anchored to your sources but stop at the answer. Chat tools write well but have never read your files. BrainBox is where both happen in one place.
How to analyze an Excel or CSV file with AI without coding
Upload the spreadsheet, ask in plain language, and the AI writes and runs the analysis code for you: tables, downloadable charts and a new Excel file with the results. You only verify.