- models
- OpenAI
- product
GPT-6 Sol vs Luna: benchmarks, pricing and which to use in BrainBox
OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026 and both are already in the BrainBox Assistant: benchmarks, Artificial Analysis score, intelligence units per task and which one to pick.
GPT-6 Sol and GPT-6 Luna have been selectable in the BrainBox Assistant since September 22, 2026, the day OpenAI released them. Both have a 1,050,000 token context window, 128,000 max output tokens and accept images. Sol sits in the Pro category of the picker and Luna in Default.
What are GPT-6 Sol and GPT-6 Luna?
They are the two tiers OpenAI added to the GPT-6 family after Astra. According to the announcement, they were trained "with similar methods as GPT-6 Astra", bringing its advances in professional work, factuality, coding, computer use and alignment "to faster, more affordable models". OpenAI states in the same text that Astra remains its best model across the board, and Astra is not in the BrainBox picker.
| GPT-6 Sol | GPT-6 Luna | |
|---|---|---|
| Built for | Complex coding and agentic workflows | Focused, high-volume tasks |
| Context | 1,050,000 tokens | 1,050,000 tokens |
| Max output | 128,000 tokens | 128,000 tokens |
| Inputs | Text and image | Text and image |
| Knowledge cutoff | April 20, 2026 | May 18, 2026 |
| Reasoning effort levels | From none to max | From none to max |
| BrainBox picker category | Pro | Default |
| Artificial Analysis Intelligence Index | 48 | 37 |
The specifications come from the official model pages for GPT-6 Sol and GPT-6 Luna.
What do the benchmarks say?
Artificial Analysis, which measures models on its own, gives GPT-6 Sol 48 points on its Intelligence Index and Luna 37. According to the leaderboard, Sol sits above Grok 4.7 and MiMo V2.6 Pro at 46, and below GPT-6 Astra and Claude Fable 5.1 at 53 and Claude Opus 5.5 at 58. Luna, at 37, is the highest score in its price range.
The benchmarks OpenAI publishes compare cost per task ahead of top score:
| Benchmark, as reported by OpenAI | GPT-6 Sol | GPT-6 Luna | Reference OpenAI cites |
|---|---|---|---|
| AutomationBench 1.0.6, business workflows | 33.2% at xhigh effort | Improves 5.4 points over its predecessor | Claude Opus 5 at max: 26.9% |
| Agents' Last Exam V1, professional work | 56.4% at max effort | Above Claude Opus 5's highest score | |
| DeepSWE v1.1, software engineering | 68.8% at max effort | 66.6% at max effort | Claude Fable 5 at xhigh: 69.9% |
| OSWorld 2.0 offline, computer use | 60.5% at xhigh effort | Claude Opus 5 at medium: 60.3% |
The best score in each row is in bold.
The competitor numbers are reported by OpenAI from public reports, not by each lab, and OpenAI notes it used Claude Fable 5 scores where Fable 5.1 scores were unavailable. On AutomationBench it also flags that the Fable 5.1 datapoint understates its cost because it omits the Opus 5 fallbacks, which happened on about 40% of tasks.
Two improvements from the announcement matter more in document work than in the table. On OpenAI's internal factuality evaluation, built from real conversations where users flagged mistakes, Sol makes "about half as many mistakes as its predecessor". And the style changed: OpenAI promises "more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance".
How does the BrainBox Assistant work with these models?
The BrainBox Assistant is an agent. You choose which Boxes and files it may use and which model does the reasoning; the agent decides the steps. Each one shows in the chat: "Listing sources", "Searching documents", "Reading document" with the pages it opened, "Saving note" when it sets something aside for later, and at the end the answer with page citations.
With a million tokens of context, either tier holds many reads in one conversation before the agent has to compress what it read into notes. The difference is how far the reasoning goes between steps, and there Sol is ahead.
Open the Assistant and choose the sources
In the sidebar you tick the Boxes and files the agent may use. You can also point at a file with @ inside the message.
Click the model picker
It sits in the bottom bar of the chat. A searchable list opens.
Type 'GPT-6'
GPT-6 Sol appears in the Pro category and GPT-6 Luna in Default, with their context and the Images badge.
Pick the tier for the task
Sol for long reviews and complex reasoning; Luna for frequent questions, summaries and batches.
Ask for the task
Page citations, the tools and the agent's visible steps work the same on both tiers.
Select model
- GPT-6 SolOpenAI
Flagship tier for complex coding and agentic workflows
1.05M contextProImages - GPT-6 LunaOpenAI
The most efficient GPT-6 tier for high-volume work
1.05M contextImages
BrainBox is an AI workspace for documents that answers with citations to the exact page. The model changes how the answer is reasoned; the trail back to the source does not.
How many intelligence units does each task use?
BrainBox does not measure usage in tokens but in intelligence units, and each task uses a different number depending on how many searches and reads the agent makes, how much the model reasons and how much it writes. Estimates for the standard plan:
| Task in the Assistant | GPT-6 Sol | GPT-6 Luna |
|---|---|---|
| Single cited question, one or two searches | ≈ 4 units | ≈ 1 unit |
| Spreadsheet analysis with the code interpreter | ≈ 11 units | ≈ 2 units |
| Case file review with about ten searches and reads | ≈ 27 units | ≈ 4 units |
| Full read of a 200 page contract, page by page | ≈ 29 units | ≈ 3 units |
| PDF report built from a case file, with code and about fifteen steps | ≈ 49 units | ≈ 5 units |
These are estimates: a longer answer, more sources or more agent steps push the number up. For reference, the Elite plan includes 400 units a month. The gap between the two tiers is why Luna is worth keeping as the everyday default: for most questions the answer is the same and the usage is a fraction.
What happens if the Box is larger than the context window?
Nothing is left out. The Assistant does not load the whole Box into the model: it searches the selected sources and reads the pages that answer the question, each with its file and page. It works the same for a Box of 200 pages or 20,000. The model decides how to reason over what it reads; the size of the Box does not decide which model you can pick.
When does Sol make sense, when Luna, and when neither?
| Task in the Assistant | Model | Why |
|---|---|---|
| Single cited question | Luna | One unit per question; the quality gap barely shows |
| Summarizing 50 meeting minutes or sorting documents in batches | Luna | Usage per question makes it viable to repeat dozens of times |
| Comparing contract versions or reviewing a case file | Sol | Reasons better between steps on long tasks |
| PDF report or data analysis with the code interpreter | Sol | OpenAI positions it for coding and agentic workflows |
| Questions about screenshots or diagrams | Either tier | Both accept images |
| An opinion where the error margin must be minimal | Claude Opus 5.5 | Ten points above Sol on the independent index |
For legal teams, the workspace model policy lets you enable Sol for the team or keep Luna only, without touching each account. Citations still point to the page with either tier, so verifying the answer works the same.
How is this different from using GPT-6 in ChatGPT?
ChatGPT Work and Codex give access to the same models, according to OpenAI. What differs is everything around them.
| ChatGPT | BrainBox with GPT-6 | |
|---|---|---|
| Sources | Whatever you paste or attach in that conversation | The Box's indexed documents: PDF, Office, transcribed audio, images |
| Citations | Do not point to a page in your files | Page, passage and expanded context for every claim |
| Switching models | OpenAI models only | Sol and Luna next to Claude, Gemini, Grok and open models, in the same chat |
| Team | Individual account | Shared Box with roles; the admin decides which models are allowed |
| Usage | Flat subscription | Intelligence units per use, based on the model you pick |
Artificial Analysis measures Sol at about 115 tokens per second and Luna at about 154 once they start writing, with over a minute of thinking before the first token at max effort. In the Assistant that shows up as a long task advancing step by step over several minutes.
Sources
- OpenAI, Introducing GPT-6 Sol and Luna: announcement of September 22, 2026, benchmarks, factuality, style and availability.
- OpenAI, model pages for GPT-6 Sol and GPT-6 Luna: context, max output, modalities, effort levels and knowledge cutoff.
- Artificial Analysis, model pages for GPT-6 Sol and GPT-6 Luna: Intelligence Index, speed and verbosity.
- Artificial Analysis, model leaderboard: scores for Claude Opus 5.5, GPT-6 Astra, Claude Fable 5.1, Grok 4.7 and MiMo V2.6 Pro.
Frequently asked questions
- What is the difference between GPT-6 Sol and GPT-6 Luna?
- Reasoning depth and usage. OpenAI describes Sol as its model for complex coding and agentic workflows, and Luna as its most efficient model for focused, high-volume tasks. Both have a 1,050,000 token context window, 128,000 max output tokens and accept images. Artificial Analysis scores Sol at 48 and Luna at 37 on its Intelligence Index. In BrainBox, a case file review is about 27 units with Sol and 4 with Luna.
- What do they cost in BrainBox?
- Usage is measured in intelligence units. Estimates on the standard plan: a single cited question, about 4 units with Sol and 1 with Luna; a data analysis with the code interpreter, about 11 with Sol and 2 with Luna; a case file review, about 27 with Sol and 4 with Luna; a PDF report, about 49 with Sol and 5 with Luna.
- Which one should I use day to day?
- Luna for most questions: one source, one search, one cited answer. Sol when the task chains many steps, such as reviewing a whole case file, cross-checking a tender against its addenda or producing a report. Switching models in the picker does not restart the chat, so you can start on Luna and move up to Sol if the answer falls short.
- Which plans include them?
- Luna sits in the Default category and Sol in the Pro category of the picker. On personal accounts, Pro models show up on paid plans; in a workspace it depends on the model policy the admin has set.
- Are they better than Claude for my documents?
- It depends on the task. On the Artificial Analysis Intelligence Index, Claude Opus 5.5 scores 58 and GPT-6 Sol 48. In the benchmarks OpenAI publishes, Sol lands within 1.1 points of Claude Fable 5 on DeepSWE v1.1 at a much lower cost per task. For an opinion where the error margin must be minimal, Opus 5.5 is still ahead; for volume and repeated tasks, Luna uses a fraction of the units.
- What if my Box is larger than the context window?
- Nothing is left out. The Assistant does not load the whole Box at once: it searches the sources you chose and reads the pages that answer the question, each with its citation. The size of the Box does not limit which model you can use.
Written by
BrainBox Team
Document intelligence, by ExaByte Company
We build BrainBox — the platform teams use to ask questions across their own documents and get answers with exact page citations.
Keep reading
Claude Opus 5.5 in BrainBox: benchmarks, pricing and when to use it
Anthropic released Claude Opus 5.5 on September 22, 2026 and it is already selectable in the BrainBox Assistant: the top spot on the Artificial Analysis Intelligence Index, with benchmarks, intelligence units per task and when it is worth it.
Grok 4.7 in BrainBox: benchmarks, pricing and when to use it
xAI's Grok 4.7 is now selectable in the BrainBox Assistant: benchmarks against Grok 4.6, its Artificial Analysis score, the intelligence units it uses per task, and which document work it suits.
Xiaomi MiMo V2.6 Pro vs Flash: benchmarks, pricing and which to use in BrainBox
Xiaomi's MiMo V2.6 Pro and Flash, MIT open weights with a 1M context window, are now in the BrainBox Assistant: benchmarks, Artificial Analysis score, intelligence units per task and which to pick.