Skip to main content
  • models
  • xAI
  • product

Grok 4.7 in BrainBox: benchmarks, pricing and when to use it

xAI's Grok 4.7 is now selectable in the BrainBox Assistant: benchmarks against Grok 4.6, its Artificial Analysis score, the intelligence units it uses per task, and which document work it suits.

BrainBox Team9 min read
View as Markdown

Grok 4.7 has been selectable in the BrainBox Assistant since September 21, 2026, the day xAI released it. It appears in the Pro category of the model picker with a 500K token window and image input. It succeeds Grok 4.6, which stays available, and in BrainBox uses fewer intelligence units on the same task.

What is Grok 4.7 and what improves over Grok 4.6?

Grok 4.7 is xAI's flagship model for coding, agents and knowledge work. According to xAI's announcement, it "uses a new, larger base model compared to Grok 4.6", was "trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete", and is "better at verifying its own work". The context window stays at 500K tokens, unchanged from Grok 4.6.

In the benchmarks xAI publishes, Grok 4.7 improves on every test over Grok 4.6, beats Claude Fable 5.1 on EEBench and on the Harvey legal benchmark, and trails it on CursorBench, Terminal-Bench and HealthBench. The competitor numbers are also xAI's, not each lab's.

Benchmark, as reported by xAIGrok 4.6Grok 4.7Claude Fable 5.1GPT-5.6 Sol
DeepSWE v1.1, software agent65.2%71.0%70.0%72.7%
CursorBench 4.0, coding in the editor40.4%46.3%51.8%41.7%
Terminal-Bench 4.0, terminal agent20.3%38.0%57.9%37.3%
HealthBench Professional48.5%56.7%62.1%60.5%
EEBench, electrical engineering53.0%64.0%56.4%39.4%
Harvey Legal Agent Benchmark15.8%19.6%6.7%2.5%

Artificial Analysis, which measures models independently, gives it 46 on its Intelligence Index v4.3, two points above Grok 4.6. That puts xAI among the four highest scoring labs, but seven points behind Claude Fable 5.1 and GPT-6 Astra, which score 53 in their maximum setting. On the same report's Coding Agent Index, Grok 4.7 with Grok Build scores 56, fourth behind Claude Fable 5.1, GPT-6 Astra and Claude Opus 5.

Grok 4.7 is talkative: according to the same report, it generates about 81,000 output tokens per index task, against 36,000 for Grok 4.6, so it reasons more and one of its answers uses more units. It also hallucinates less: on AA-Omniscience, the Artificial Analysis knowledge test, its hallucination rate drops from 34% to 29%, and accuracy almost unchanged, 47% against 48%.

How does the BrainBox Assistant work with Grok 4.7?

The BrainBox Assistant is an agent. You choose which Boxes and files it may use and which model does the reasoning; the agent decides the steps. Each one shows in the chat: "Listing sources", "Searching documents", "Reading document" with the pages it opened, "Saving note" when it sets something aside for later, and at the end the answer with page citations. For a large case file it does not load everything at once: it searches, reads the pages that matter, and if the task calls for it, reads a whole document.

Grok 4.7 was trained for tasks that chain many steps, and that is what the Assistant does in a review: several searches, several reads, a cross-check between documents and a long answer. That is where the gain over Grok 4.6 shows.

How do you pick Grok 4.7 in BrainBox?

  1. Open the Assistant and choose the sources

    In the sidebar you tick the Boxes and files the agent may use. You can also point at a file with @ inside the message.

  2. Click the model picker

    It sits in the bottom bar of the chat. A searchable list opens.

  3. Type 'Grok'

    Grok 4.7 and Grok 4.6 appear with their context and the Pro and Images badges.

  4. Choose Grok 4.7

    If it is missing, your personal plan does not include the Pro category or the workspace policy does not allow it.

  5. Ask for the task

    Page citations, the tools and the agent's visible steps work the same with any model.

Select model

Select model

Grok
  • xAI logo
    Grok 4.7xAI

    Long-running coding, agents and knowledge work

    500K contextProImages
  • xAI logo
    Grok 4.6xAI

    Coding, knowledge work and STEM; succeeded by Grok 4.7

    500K contextProImages
Simulation of the BrainBox Assistant model picker with Grok 4.7 selected.

BrainBox is an AI workspace for documents that answers with citations to the exact page. The model changes how the answer is reasoned; the trail back to the source does not.

How many intelligence units does Grok 4.7 use per task?

BrainBox does not measure usage in tokens but in intelligence units, and each task uses a different number depending on how many searches and reads the agent makes, how much the model reasons and how much it writes. Estimates for the standard plan:

Task in the AssistantGrok 4.7Grok 4.6
Single cited question, one or two searches≈ 3 units≈ 3 units
Spreadsheet analysis with the code interpreter≈ 8 units≈ 9 units
Case file review with about ten searches and reads≈ 20 units≈ 25 units
Full read of a 200 page contract, page by page≈ 22 units≈ 28 units
PDF report built from a case file, with code and about fifteen steps≈ 35 units≈ 42 units
A 450 page case file read whole in a single pass≈ 95 units≈ 120 units

These are estimates: a longer answer, more selected sources or more agent steps push the number up. For reference, the Elite plan includes 400 units a month.

What happens if the Box is larger than Grok 4.7's context?

Nothing is left out. The Assistant does not load the whole Box into the model: it searches the selected sources and reads the pages that answer the question, each with its file and page. It works the same for a Box of 200 pages or 20,000. The model decides how to reason over what it reads; the size of the Box does not decide which model you can pick.

Which tasks is it worth choosing for?

Grok 4.7 fits multi-step tasks: reviewing a case file of dozens of documents, comparing two versions of a contract, cross-checking a tender against its addenda, or asking for a report delivered as a PDF. On the Harvey Legal Agent Benchmark, as reported by xAI, Grok 4.7 scores 19.6% against 15.8% for Grok 4.6, 6.7% for Claude Fable 5.1 and 2.5% for GPT-5.6 Sol. That is a low number in absolute terms for every model. A legal memo drafted with Grok 4.7 gets reviewed the same way as one drafted with any other model.

Task in the AssistantGrok 4.7?Alternative
Single cited questionNot needed; a default model answers the same with fewer unitsDefault or Flash
Full case file reviewYes, long multi-step work is its strengthClaude or GPT if budget allows
Data analysis with the code interpreterYes, better at coding than 4.6Claude Fable 5.1 leads on coding agents
Questions about screenshots or diagramsYes, it accepts imagesAny other model with the Images badge
Reading more than 300 pages in a single passCarefully, unit usage nearly doublesXiaomi MiMo V2.6, 1M context with no usage cliff

You can test it inside the same conversation: switch from Grok 4.6 to 4.7 in the picker, repeat the question and compare. Citations still point to the page, so verifying the answer works the same with both.

How is this different from using Grok in xAI's app?

The Grok app, Cursor and Grok Build give access to the same model, according to xAI. What differs is everything around it.

Grok appBrainBox with Grok 4.7
SourcesWhatever you paste or attach in that conversationThe Box's indexed documents: PDF, Office, transcribed audio, images
CitationsDo not point to a page in your filesPage, passage and expanded context for every claim
Switching modelsxAI models onlyGrok 4.7 next to Claude, GPT, Gemini and open models, in the same chat
TeamIndividual accountShared Box with roles; the admin decides which models are allowed
UsageFlat subscriptionIntelligence units per use, based on the model you pick

For legal teams, the workspace model policy lets you enable Grok 4.7 for the team or block it, without touching each account. The rest of the document workspace works the same with any model.

A full task on the Artificial Analysis index took Grok 4.7 about 7.1 minutes on average; in BrainBox, a long review can take several minutes and you watch it advance step by step in the chat.

Sources

  • xAI, Introducing Grok 4.7: announcement of September 21, 2026, benchmark table and availability.
  • Artificial Analysis, Benchmarking Grok 4.7: Intelligence Index, Coding Agent Index, tokens per task, AA-Omniscience, 500K context.
  • Artificial Analysis, model leaderboard: scores for Claude Fable 5.1, GPT-6 Astra and Grok 4.6.
  • Artificial Analysis, Grok 4.7 model page: context and modalities.

Frequently asked questions

What changed between Grok 4.6 and Grok 4.7?
According to xAI, Grok 4.7 uses a larger base model, a longer reinforcement learning run on tasks that take hours, and checks its own output more carefully. On DeepSWE v1.1 it goes from 65.2% to 71.0%; on the Harvey Legal Agent Benchmark, from 15.8% to 19.6%. On the Artificial Analysis Intelligence Index it moves from 44 to 46. The context window stays at 500K tokens. In BrainBox the same task also uses 10% to 20% fewer intelligence units than with Grok 4.6.
How much does Grok 4.7 cost in BrainBox?
Usage is measured in intelligence units, not tokens. Estimates on the standard plan: a single cited question, about 3 units; a data analysis with the code interpreter, about 8; a case file review with searches and page reads, about 20; a PDF report built from a case file, about 35. Each uses 10% to 20% fewer units than Grok 4.6.
Which plans include it?
Grok 4.7 sits in the Pro category of the picker. On personal accounts it shows up on paid plans; in a workspace it depends on the model policy the admin has set. If you cannot see it, check the workspace AI policy.
Does Grok 4.7 read images?
Yes. In BrainBox it is flagged with image support, so you can attach a screenshot or a diagram to a question, or ask about the images indexed in the Box.
Is it better than Claude or GPT for my documents?
On the Artificial Analysis Intelligence Index, Claude Fable 5.1 and GPT-6 Astra score 53 against 46 for Grok 4.7. It also trails on coding agents. What it has going for it is usage: a long case file review uses noticeably fewer intelligence units than with those models. On a single cited question, neither the quality gap nor the usage gap shows.
What if my Box is larger than Grok 4.7's context window?
Nothing is left out. The Assistant does not load the whole Box at once: it searches the sources you chose and reads the pages that answer the question, each with its citation. The size of the Box does not limit which model you can use.
Is Grok 4.6 still available?
Yes. Grok 4.6 stays in the picker and uses somewhat more intelligence units per task. Chats that were already using 4.6 do not switch models on their own.

Written by

BrainBox Team

Document intelligence, by ExaByte Company

We build BrainBox — the platform teams use to ask questions across their own documents and get answers with exact page citations.