# Grok 4.7 in BrainBox: benchmarks, pricing and when to use it

> xAI's Grok 4.7 is now selectable in the BrainBox Assistant: benchmarks against Grok 4.6, its Artificial Analysis score, the intelligence units it uses per task, and which document work it suits.

- Source: https://www.usebrainbox.com/en/blog/grok-4-7-in-brainbox-benchmarks-pricing-and-when-to-use-it
- Author: BrainBox Team
- Published: 2026-09-21
- Updated: 2026-09-22
- Language: en
- Tags: models, xAI, product

## Summary

Grok 4.7 is xAI's new flagship model, released on September 21, 2026, and available in the BrainBox Assistant from that day with a 500K token context window and image input. Artificial Analysis scores it 46 on its Intelligence Index v4.3, two points above Grok 4.6, and xAI reports 71.0% on DeepSWE v1.1. In BrainBox it uses 10% to 20% fewer intelligence units than Grok 4.6 on the same task: about 3 per cited question and about 20 per case file review. It suits long multi-step jobs; for a short question, the default model uses less.

---
Grok 4.7 has been selectable in the BrainBox Assistant since September 21, 2026, the day [xAI released it](https://x.ai/news/grok-4-7). It appears in the Pro category of the model picker with a [500K token](https://artificialanalysis.ai/articles/benchmarking-grok-4-7) window and image input. It succeeds Grok 4.6, which stays available, and in BrainBox uses fewer intelligence units on the same task.

> **Remember**
>
> Grok 4.7 improves on long coding and agent tasks and uses fewer intelligence units than Grok 4.6. In BrainBox it is worth choosing for case file reviews with many steps; for a short cited question, the default model still uses less.

## What is Grok 4.7 and what improves over Grok 4.6?

Grok 4.7 is xAI's flagship model for coding, agents and knowledge work. According to [xAI's announcement](https://x.ai/news/grok-4-7), it "uses a new, larger base model compared to Grok 4.6", was "trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete", and is "better at verifying its own work". The context window stays at 500K tokens, [unchanged from Grok 4.6](https://artificialanalysis.ai/articles/benchmarking-grok-4-7).

In the benchmarks [xAI publishes](https://x.ai/news/grok-4-7), Grok 4.7 improves on every test over Grok 4.6, beats Claude Fable 5.1 on EEBench and on the Harvey legal benchmark, and trails it on CursorBench, Terminal-Bench and HealthBench. The competitor numbers are also xAI's, not each lab's.

| Benchmark, as reported by xAI | Grok 4.6 | Grok 4.7 | Claude Fable 5.1 | GPT-5.6 Sol |
|---|---|---|---|---|
| DeepSWE v1.1, software agent | 65.2% | 71.0% | 70.0% | 72.7% |
| CursorBench 4.0, coding in the editor | 40.4% | 46.3% | 51.8% | 41.7% |
| Terminal-Bench 4.0, terminal agent | 20.3% | 38.0% | 57.9% | 37.3% |
| HealthBench Professional | 48.5% | 56.7% | 62.1% | 60.5% |
| EEBench, electrical engineering | 53.0% | 64.0% | 56.4% | 39.4% |
| Harvey Legal Agent Benchmark | 15.8% | 19.6% | 6.7% | 2.5% |

[Artificial Analysis](https://artificialanalysis.ai/articles/benchmarking-grok-4-7), which measures models independently, gives it 46 on its Intelligence Index v4.3, two points above Grok 4.6. That puts xAI among the four highest scoring labs, but seven points behind Claude Fable 5.1 and GPT-6 Astra, which [score 53](https://artificialanalysis.ai/leaderboards/models) in their maximum setting. On the same report's Coding Agent Index, Grok 4.7 with Grok Build scores 56, fourth behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5.

Grok 4.7 is talkative: according to [the same report](https://artificialanalysis.ai/articles/benchmarking-grok-4-7), it generates about 81,000 output tokens per index task, against 36,000 for Grok 4.6, so it reasons more and one of its answers uses more units. It also hallucinates less: on AA-Omniscience, the Artificial Analysis knowledge test, its hallucination rate drops from 34% to 29%, and accuracy almost unchanged, 47% against 48%.

## How does the BrainBox Assistant work with Grok 4.7?

The BrainBox Assistant is an agent. You choose which Boxes and files it may use and which model does the reasoning; the agent decides the steps. Each one shows in the chat: "Listing sources", "Searching documents", "Reading document" with the pages it opened, "Saving note" when it sets something aside for later, and at the end the answer with page citations. For a large case file it does not load everything at once: it searches, reads the pages that matter, and if the task calls for it, reads a whole document.

Grok 4.7 was trained for tasks that chain many steps, and that is what the Assistant does in a review: several searches, several reads, a cross-check between documents and a long answer. That is where the gain over Grok 4.6 shows.

## How do you pick Grok 4.7 in BrainBox?

  - **Open the Assistant and choose the sources** In the sidebar you tick the Boxes and files the agent may use. You can also point at a file with @ inside the message.
  - **Click the model picker** It sits in the bottom bar of the chat. A searchable list opens.
  - **Type 'Grok'** Grok 4.7 and Grok 4.6 appear with their context and the Pro and Images badges.
  - **Choose Grok 4.7** If it is missing, your personal plan does not include the Pro category or the workspace policy does not allow it.
  - **Ask for the task** Page citations, the tools and the agent's visible steps work the same with any model.

*Simulation of the BrainBox Assistant model picker with Grok 4.7 selected.*

BrainBox is an AI workspace for documents that answers with citations to the exact page. The model changes how the answer is reasoned; the trail back to the source does not.

## How many intelligence units does Grok 4.7 use per task?

BrainBox does not measure usage in tokens but in intelligence units, and each task uses a different number depending on how many searches and reads the agent makes, how much the model reasons and how much it writes. Estimates for the standard plan:

| Task in the Assistant | Grok 4.7 | Grok 4.6 |
|---|---|---|
| Single cited question, one or two searches | ≈ 3 units | ≈ 3 units |
| Spreadsheet analysis with the code interpreter | ≈ 8 units | ≈ 9 units |
| Case file review with about ten searches and reads | ≈ 20 units | ≈ 25 units |
| Full read of a 200 page contract, page by page | ≈ 22 units | ≈ 28 units |
| PDF report built from a case file, with code and about fifteen steps | ≈ 35 units | ≈ 42 units |
| A 450 page case file read whole in a single pass | ≈ 95 units | ≈ 120 units |

These are estimates: a longer answer, more selected sources or more agent steps push the number up. For reference, the Elite plan includes 400 units a month.

> **The 300 page cliff**
>
> With Grok 4.7, a single step carrying more than about 300 pages of context uses close to twice the units per page. The Assistant rarely gets there because it reads by pages and keeps notes, but if you ask it to "read the whole case file in one go" on 450 pages, expect about 95 units against the 20 of a review in parts. For large case files, ask for the review and let the agent choose what to read.

## What happens if the Box is larger than Grok 4.7's context?

Nothing is left out. The Assistant does not load the whole Box into the model: it searches the selected sources and reads the pages that answer the question, each with its file and page. It works the same for a Box of 200 pages or 20,000. The model decides how to reason over what it reads; the size of the Box does not decide which model you can pick.

## Which tasks is it worth choosing for?

Grok 4.7 fits multi-step tasks: reviewing a case file of dozens of documents, [comparing two versions of a contract](/en/blog/compare-two-contract-versions-with-ai), cross-checking a tender against its addenda, or [asking for a report delivered as a PDF](/en/blog/ask-ai-for-a-report-and-get-a-designed-pdf-or-powerpoint). On the Harvey Legal Agent Benchmark, [as reported by xAI](https://x.ai/news/grok-4-7), Grok 4.7 scores 19.6% against 15.8% for Grok 4.6, 6.7% for Claude Fable 5.1 and 2.5% for GPT-5.6 Sol. That is a low number in absolute terms for every model. A legal memo drafted with Grok 4.7 gets reviewed the same way as one drafted with any other model.

| Task in the Assistant | Grok 4.7? | Alternative |
|---|---|---|
| Single cited question | Not needed; a default model answers the same with fewer units | Default or Flash |
| Full case file review | Yes, long multi-step work is its strength | Claude or GPT if budget allows |
| Data analysis with the code interpreter | Yes, better at coding than 4.6 | Claude Fable 5.1 [leads](https://artificialanalysis.ai/articles/benchmarking-grok-4-7) on coding agents |
| Questions about screenshots or diagrams | Yes, it accepts images | Any other model with the Images badge |
| Reading more than 300 pages in a single pass | Carefully, unit usage nearly doubles | [Xiaomi MiMo V2.6](/en/blog/xiaomi-mimo-v2-6-pro-vs-flash-benchmarks-pricing-which-to-use-in-brainbox), 1M context with no usage cliff |

You can test it inside the same conversation: switch from Grok 4.6 to 4.7 in the picker, repeat the question and compare. Citations still point to the page, so [verifying the answer](/en/blog/how-to-verify-an-ai-answer-about-your-documents) works the same with both.

## How is this different from using Grok in xAI's app?

The Grok app, Cursor and Grok Build give access to the same model, [according to xAI](https://x.ai/news/grok-4-7). What differs is everything around it.

| | Grok app | BrainBox with Grok 4.7 |
|---|---|---|
| Sources | Whatever you paste or attach in that conversation | The Box's indexed documents: PDF, Office, transcribed audio, images |
| Citations | Do not point to a page in your files | Page, passage and expanded context for every claim |
| Switching models | xAI models only | Grok 4.7 next to Claude, GPT, Gemini and open models, in the same chat |
| Team | Individual account | Shared Box with roles; the admin decides which models are allowed |
| Usage | Flat subscription | Intelligence units per use, based on the model you pick |

For legal teams, the workspace model policy lets you enable Grok 4.7 for the team or block it, without touching each account. The rest of the [document workspace](/en/blog/brainbox-2026-workspace-for-document-work) works the same with any model.

A full task on the Artificial Analysis index took Grok 4.7 [about 7.1 minutes on average](https://artificialanalysis.ai/articles/benchmarking-grok-4-7); in BrainBox, a long review can take several minutes and you watch it advance step by step in the chat.

## Sources

- xAI, [Introducing Grok 4.7](https://x.ai/news/grok-4-7): announcement of September 21, 2026, benchmark table and availability.
- Artificial Analysis, [Benchmarking Grok 4.7](https://artificialanalysis.ai/articles/benchmarking-grok-4-7): Intelligence Index, Coding Agent Index, tokens per task, AA-Omniscience, 500K context.
- Artificial Analysis, [model leaderboard](https://artificialanalysis.ai/leaderboards/models): scores for Claude Fable 5.1, GPT-6 Astra and Grok 4.6.
- Artificial Analysis, [Grok 4.7 model page](https://artificialanalysis.ai/models/grok-4-7): context and modalities.

## FAQ

### What changed between Grok 4.6 and Grok 4.7?

According to xAI, Grok 4.7 uses a larger base model, a longer reinforcement learning run on tasks that take hours, and checks its own output more carefully. On DeepSWE v1.1 it goes from 65.2% to 71.0%; on the Harvey Legal Agent Benchmark, from 15.8% to 19.6%. On the Artificial Analysis Intelligence Index it moves from 44 to 46. The context window stays at 500K tokens. In BrainBox the same task also uses 10% to 20% fewer intelligence units than with Grok 4.6.

### How much does Grok 4.7 cost in BrainBox?

Usage is measured in intelligence units, not tokens. Estimates on the standard plan: a single cited question, about 3 units; a data analysis with the code interpreter, about 8; a case file review with searches and page reads, about 20; a PDF report built from a case file, about 35. Each uses 10% to 20% fewer units than Grok 4.6.

### Which plans include it?

Grok 4.7 sits in the Pro category of the picker. On personal accounts it shows up on paid plans; in a workspace it depends on the model policy the admin has set. If you cannot see it, check the workspace AI policy.

### Does Grok 4.7 read images?

Yes. In BrainBox it is flagged with image support, so you can attach a screenshot or a diagram to a question, or ask about the images indexed in the Box.

### Is it better than Claude or GPT for my documents?

On the Artificial Analysis Intelligence Index, Claude Fable 5.1 and GPT-6 Astra score 53 against 46 for Grok 4.7. It also trails on coding agents. What it has going for it is usage: a long case file review uses noticeably fewer intelligence units than with those models. On a single cited question, neither the quality gap nor the usage gap shows.

### What if my Box is larger than Grok 4.7's context window?

Nothing is left out. The Assistant does not load the whole Box at once: it searches the sources you chose and reads the pages that answer the question, each with its citation. The size of the Box does not limit which model you can use.

### Is Grok 4.6 still available?

Yes. Grok 4.6 stays in the picker and uses somewhat more intelligence units per task. Chats that were already using 4.6 do not switch models on their own.
