# GPT-6 Sol vs Luna: benchmarks, pricing and which to use in BrainBox

> OpenAI released GPT-6 Sol and GPT-6 Luna on September 22, 2026 and both are already in the BrainBox Assistant: benchmarks, Artificial Analysis score, intelligence units per task and which one to pick.

- Source: https://www.usebrainbox.com/en/blog/gpt-6-sol-vs-luna-benchmarks-pricing-which-to-use-in-brainbox
- Author: BrainBox Team
- Published: 2026-09-22
- Language: en
- Tags: models, OpenAI, product

## Summary

GPT-6 Sol and GPT-6 Luna are the models OpenAI released on September 22, 2026, and both can be selected in the BrainBox Assistant with a 1,050,000 token context window and image input. Sol is the tier for complex work and agents; Luna is the efficient tier for frequent questions and batches. Artificial Analysis scores them 48 and 37 on its Intelligence Index. A case file review is about 27 intelligence units with Sol and 4 with Luna.

---
GPT-6 Sol and GPT-6 Luna have been selectable in the BrainBox Assistant since September 22, 2026, the day [OpenAI released them](https://openai.com/index/introducing-gpt-6-sol-and-luna/). Both have a [1,050,000 token context window](https://developers.openai.com/api/docs/models/gpt-6-sol), 128,000 max output tokens and accept images. Sol sits in the Pro category of the picker and Luna in Default.

> **Remember**
>
> Sol and Luna are the same quality jump at two usage tiers. A case file review is about 27 intelligence units with Sol and 4 with Luna. If in doubt, start on Luna.

## What are GPT-6 Sol and GPT-6 Luna?

They are the two tiers OpenAI added to the GPT-6 family after Astra. According to [the announcement](https://openai.com/index/introducing-gpt-6-sol-and-luna/), they were trained "with similar methods as GPT-6 Astra", bringing its advances in professional work, factuality, coding, computer use and alignment "to faster, more affordable models". OpenAI states in the same text that Astra remains its best model across the board, and Astra is not in the BrainBox picker.

| | GPT-6 Sol | GPT-6 Luna |
|---|---|---|
| Built for | Complex coding and agentic workflows | Focused, high-volume tasks |
| Context | 1,050,000 tokens | 1,050,000 tokens |
| Max output | 128,000 tokens | 128,000 tokens |
| Inputs | Text and image | Text and image |
| Knowledge cutoff | April 20, 2026 | May 18, 2026 |
| Reasoning effort levels | From none to max | From none to max |
| BrainBox picker category | Pro | Default |
| Artificial Analysis Intelligence Index | 48 | 37 |

The specifications come from the official model pages for [GPT-6 Sol](https://developers.openai.com/api/docs/models/gpt-6-sol) and [GPT-6 Luna](https://developers.openai.com/api/docs/models/gpt-6-luna).

## What do the benchmarks say?

[Artificial Analysis](https://artificialanalysis.ai/models/gpt-6-sol), which measures models on its own, gives GPT-6 Sol 48 points on its Intelligence Index and [Luna 37](https://artificialanalysis.ai/models/gpt-6-luna). According to [the leaderboard](https://artificialanalysis.ai/leaderboards/models), Sol sits above Grok 4.7 and MiMo V2.6 Pro at 46, and below GPT-6 Astra and Claude Fable 5.1 at 53 and Claude Opus 5.5 at 58. Luna, at 37, is the highest score in its price range.

The benchmarks [OpenAI publishes](https://openai.com/index/introducing-gpt-6-sol-and-luna/) compare cost per task ahead of top score:

| Benchmark, as reported by OpenAI | GPT-6 Sol | GPT-6 Luna | Reference OpenAI cites |
| --- | --- | --- | --- |
| AutomationBench 1.0.6, business workflows | **33.2%** at xhigh effort | Improves 5.4 points over its predecessor | Claude Opus 5 at max: 26.9% |
| Agents' Last Exam V1, professional work | **56.4%** at max effort | | Above Claude Opus 5's highest score |
| DeepSWE v1.1, software engineering | 68.8% at max effort | 66.6% at max effort | Claude Fable 5 at xhigh: **69.9%** |
| OSWorld 2.0 offline, computer use | **60.5%** at xhigh effort | | Claude Opus 5 at medium: 60.3% |

The best score in each row is in bold.

The competitor numbers are reported by OpenAI from public reports, not by each lab, and OpenAI notes it used Claude Fable 5 scores where Fable 5.1 scores were unavailable. On AutomationBench it also flags that the Fable 5.1 datapoint understates its cost because it omits the Opus 5 fallbacks, which happened on about 40% of tasks.

Two improvements from the announcement matter more in document work than in the table. On OpenAI's internal factuality evaluation, built from real conversations where users flagged mistakes, Sol makes "about half as many mistakes as its predecessor". And the style changed: OpenAI promises "more clarity, less jargon, fewer odd turns of phrase, fewer low-value details, and slightly shorter answers overall without losing substance".

## How does the BrainBox Assistant work with these models?

The BrainBox Assistant is an agent. You choose which Boxes and files it may use and which model does the reasoning; the agent decides the steps. Each one shows in the chat: "Listing sources", "Searching documents", "Reading document" with the pages it opened, "Saving note" when it sets something aside for later, and at the end the answer with page citations.

With a million tokens of context, either tier holds many reads in one conversation before the agent has to compress what it read into notes. The difference is how far the reasoning goes between steps, and there Sol is ahead.

  - **Open the Assistant and choose the sources** In the sidebar you tick the Boxes and files the agent may use. You can also point at a file with @ inside the message.
  - **Click the model picker** It sits in the bottom bar of the chat. A searchable list opens.
  - **Type 'GPT-6'** GPT-6 Sol appears in the Pro category and GPT-6 Luna in Default, with their context and the Images badge.
  - **Pick the tier for the task** Sol for long reviews and complex reasoning; Luna for frequent questions, summaries and batches.
  - **Ask for the task** Page citations, the tools and the agent's visible steps work the same on both tiers.

*Simulation of the BrainBox Assistant model picker with the two GPT-6 tiers.*

BrainBox is an AI workspace for documents that answers with citations to the exact page. The model changes how the answer is reasoned; the trail back to the source does not.

## How many intelligence units does each task use?

BrainBox does not measure usage in tokens but in intelligence units, and each task uses a different number depending on how many searches and reads the agent makes, how much the model reasons and how much it writes. Estimates for the standard plan:

| Task in the Assistant | GPT-6 Sol | GPT-6 Luna |
|---|---|---|
| Single cited question, one or two searches | ≈ 4 units | ≈ 1 unit |
| Spreadsheet analysis with the code interpreter | ≈ 11 units | ≈ 2 units |
| Case file review with about ten searches and reads | ≈ 27 units | ≈ 4 units |
| Full read of a 200 page contract, page by page | ≈ 29 units | ≈ 3 units |
| PDF report built from a case file, with code and about fifteen steps | ≈ 49 units | ≈ 5 units |

These are estimates: a longer answer, more sources or more agent steps push the number up. For reference, the Elite plan includes 400 units a month. The gap between the two tiers is why Luna is worth keeping as the everyday default: for most questions the answer is the same and the usage is a fraction.

> **Start on Luna and move up to Sol if you need to**
>
> Switching models in the picker does not restart the chat. Ask with Luna, and if the answer falls short or the agent loses the thread between steps, switch to Sol and ask again. If both cite the same pages and say the same thing, Luna is enough for that task.

## What happens if the Box is larger than the context window?

Nothing is left out. The Assistant does not load the whole Box into the model: it searches the selected sources and reads the pages that answer the question, each with its file and page. It works the same for a Box of 200 pages or 20,000. The model decides how to reason over what it reads; the size of the Box does not decide which model you can pick.

## When does Sol make sense, when Luna, and when neither?

| Task in the Assistant | Model | Why |
|---|---|---|
| Single cited question | Luna | One unit per question; the quality gap barely shows |
| Summarizing 50 meeting minutes or sorting documents in batches | Luna | Usage per question makes it viable to repeat dozens of times |
| [Comparing contract versions](/en/blog/compare-two-contract-versions-with-ai) or reviewing a case file | Sol | Reasons better between steps on long tasks |
| [PDF report](/en/blog/ask-ai-for-a-report-and-get-a-designed-pdf-or-powerpoint) or data analysis with the code interpreter | Sol | OpenAI positions it for coding and agentic workflows |
| Questions about screenshots or diagrams | Either tier | Both accept images |
| An opinion where the error margin must be minimal | [Claude Opus 5.5](/en/blog/claude-opus-5-5-in-brainbox-benchmarks-pricing-and-when-to-use-it) | Ten points above Sol on the independent index |

For [legal teams](/en/solutions/legal), the workspace model policy lets you enable Sol for the team or keep Luna only, without touching each account. Citations still point to the page with either tier, so [verifying the answer](/en/blog/how-to-verify-an-ai-answer-about-your-documents) works the same.

## How is this different from using GPT-6 in ChatGPT?

ChatGPT Work and Codex give access to the same models, [according to OpenAI](https://openai.com/index/introducing-gpt-6-sol-and-luna/). What differs is everything around them.

| | ChatGPT | BrainBox with GPT-6 |
|---|---|---|
| Sources | Whatever you paste or attach in that conversation | The Box's indexed documents: PDF, Office, transcribed audio, images |
| Citations | Do not point to a page in your files | Page, passage and expanded context for every claim |
| Switching models | OpenAI models only | Sol and Luna next to Claude, Gemini, Grok and open models, in the same chat |
| Team | Individual account | Shared Box with roles; the admin decides which models are allowed |
| Usage | Flat subscription | Intelligence units per use, based on the model you pick |

Artificial Analysis measures Sol at about 115 tokens per second and [Luna at about 154](https://artificialanalysis.ai/models/gpt-6-luna) once they start writing, with over a minute of thinking before the first token at max effort. In the Assistant that shows up as a long task advancing step by step over several minutes.

## Sources

- OpenAI, [Introducing GPT-6 Sol and Luna](https://openai.com/index/introducing-gpt-6-sol-and-luna/): announcement of September 22, 2026, benchmarks, factuality, style and availability.
- OpenAI, model pages for [GPT-6 Sol](https://developers.openai.com/api/docs/models/gpt-6-sol) and [GPT-6 Luna](https://developers.openai.com/api/docs/models/gpt-6-luna): context, max output, modalities, effort levels and knowledge cutoff.
- Artificial Analysis, model pages for [GPT-6 Sol](https://artificialanalysis.ai/models/gpt-6-sol) and [GPT-6 Luna](https://artificialanalysis.ai/models/gpt-6-luna): Intelligence Index, speed and verbosity.
- Artificial Analysis, [model leaderboard](https://artificialanalysis.ai/leaderboards/models): scores for Claude Opus 5.5, GPT-6 Astra, Claude Fable 5.1, Grok 4.7 and MiMo V2.6 Pro.

## FAQ

### What is the difference between GPT-6 Sol and GPT-6 Luna?

Reasoning depth and usage. OpenAI describes Sol as its model for complex coding and agentic workflows, and Luna as its most efficient model for focused, high-volume tasks. Both have a 1,050,000 token context window, 128,000 max output tokens and accept images. Artificial Analysis scores Sol at 48 and Luna at 37 on its Intelligence Index. In BrainBox, a case file review is about 27 units with Sol and 4 with Luna.

### What do they cost in BrainBox?

Usage is measured in intelligence units. Estimates on the standard plan: a single cited question, about 4 units with Sol and 1 with Luna; a data analysis with the code interpreter, about 11 with Sol and 2 with Luna; a case file review, about 27 with Sol and 4 with Luna; a PDF report, about 49 with Sol and 5 with Luna.

### Which one should I use day to day?

Luna for most questions: one source, one search, one cited answer. Sol when the task chains many steps, such as reviewing a whole case file, cross-checking a tender against its addenda or producing a report. Switching models in the picker does not restart the chat, so you can start on Luna and move up to Sol if the answer falls short.

### Which plans include them?

Luna sits in the Default category and Sol in the Pro category of the picker. On personal accounts, Pro models show up on paid plans; in a workspace it depends on the model policy the admin has set.

### Are they better than Claude for my documents?

It depends on the task. On the Artificial Analysis Intelligence Index, Claude Opus 5.5 scores 58 and GPT-6 Sol 48. In the benchmarks OpenAI publishes, Sol lands within 1.1 points of Claude Fable 5 on DeepSWE v1.1 at a much lower cost per task. For an opinion where the error margin must be minimal, Opus 5.5 is still ahead; for volume and repeated tasks, Luna uses a fraction of the units.

### What if my Box is larger than the context window?

Nothing is left out. The Assistant does not load the whole Box at once: it searches the sources you chose and reads the pages that answer the question, each with its citation. The size of the Box does not limit which model you can use.
