
GPT-6 Sol is the middle model in OpenAI's GPT-6 family, released on September 22, 2026. It sits between Luna, the cheap volume tier, and Astra, the flagship.
It costs $2 per million input tokens and $10 per million output tokens. That is half the price of GPT-5.6 Sol and one fifth of Astra.
OpenAI built it "to power complex coding and agentic workflows." The short answer for most teams: use Sol as the default for work that needs judgment, and treat Astra as the exception you have to justify.
GPT-6 Sol at a Glance
| GPT-6 Sol | |
|---|---|
| Position | Middle GPT-6 tier, above GPT-6 Luna and below GPT-6 Astra; replaces GPT-5.6 Sol |
| Best for | Coding agents, multi-step business workflows, research and synthesis |
| Input price | $2.00 per 1M tokens ($0.20 cached) |
| Output price | $10.00 per 1M tokens |
| Context window | 1,050,000 tokens, up to 128,000 output tokens |
| Long-prompt rule | Above 272K input tokens, the whole request bills at 2x input and 1.5x output |
| Reasoning effort | none, low, medium (default), high, xhigh, max |
| Tools | Function calling, structured outputs, web search, computer use, hosted shell, apply patch, MCP (Responses API) |
| Model ID | gpt-6-sol on OpenAI, openai/gpt-6-sol on Modellix |
Specs and prices from OpenAI's model page.
GPT-6 Sol vs GPT-5.6 Sol
| GPT-5.6 Sol | GPT-6 Sol | |
|---|---|---|
| Input / output per 1M | $4 / $20 | $2 / $10 |
| FrontierCode 1.1 Main | Baseline | "Improves substantially," matching Fable 5.1 (xhigh) |
| Factuality (OpenAI internal) | Baseline | About half as many mistakes |
| Answer style | Longer, restates details | Shorter, says what it checked |
The price cut applies to every token, so an unchanged workload costs half as much.
One number does not move in Sol's favor. OpenAI's Astra launch table lists GPT-5.6 Sol at 72.7% on DeepSWE 1.1, while the Sol post reports GPT-6 Sol at 68.8% at max effort. The two posts may not use the same setup, but it is a reason to test coding-heavy workloads before switching blind.
If you run GPT-5.6 Sol today, switching is mostly a model-ID change. Rerun your own evaluation anyway, because shorter answers can break parsers that expect the old format.
GPT-6 Sol Pricing
| Rate (per 1M tokens) | ≤272K input | >272K input |
|---|---|---|
| Input | $2.00 | $4.00 |
| Cached input read | $0.20 | $0.40 |
| Cache write | $2.50 | $5.00 |
| Output | $10.00 | $15.00 |
Batch and Flex are half price on OpenAI. Fast mode is double.
The same workload on all three GPT-6 models
| Workload | Luna | Sol | Astra |
|---|---|---|---|
| 10M input + 2M output a month | $2.00 | $40.00 | $200.00 |
| Same, 80% of input served from cache | $1.28 | $25.60 | $128.00 |
| One 400K-token prompt, 5K output | $0.084 | $1.68 | $8.38 |
These are our calculations from OpenAI list prices, excluding cache-write charges.
Caching matters more on Sol than on Luna in absolute terms. Agents that resend the same system prompt and tool definitions every turn can cut a third off the bill at an 80% hit rate, and OpenAI says GPT-6 caching now holds across changes in reasoning effort and tool availability.
For Astra's full rate card, including cache and long-context tiers, see our GPT-6 Astra pricing breakdown.
GPT-6 Sol Benchmarks
All numbers below are from OpenAI's launch post.
| Benchmark | What it measures | GPT-6 Sol | Compared with | Result | Cost per task |
|---|---|---|---|---|---|
| AutomationBench 1.0.6 | End-to-end business workflows across 47 tools | 33.2% xhigh | GPT-6 Astra (low) 30.3%, Claude Opus 5 (max) 26.9% | Beats both | $0.27; Astra costs 3.9x, Opus 5 11.1x |
| Agents' Last Exam | Long professional tasks across 55 sub-industries | 56.4% max | Claude Opus 5's best score | Beats Opus 5 | 60% lower than Opus 5 |
| OSWorld 2.0 offline | Computer-use workflows, partial reward | 60.5% xhigh | Claude Opus 5 (medium) 60.3% | Ties | About 80% lower |
| FrontierCode 1.1 Main | Whether code is ready to merge, not just correct | Not published | Claude Fable 5.1 (xhigh) | Matches | "Much lower" (not quantified) |
| DeepSWE 1.1 | Long software engineering tasks in real codebases | 68.8% max | Claude Fable 5 best, 69.9% (xhigh) | 1.1 pts behind | About 80% lower |
Two things to keep in mind when reading this table.
First, every Sol score is at xhigh or max effort. Those settings produce more reasoning tokens, and reasoning tokens bill as output.
Second, the comparisons are cost-per-task claims as much as score claims. Sol's case is "close to the frontier for much less," not "best in class."
What Can You Build with GPT-6 Sol
Sol fits products where the model has to reason over several steps and the output goes to a customer or a codebase.
- Pull request reviewer. A bot that reads a diff, flags breaking changes, and explains why, before a human approves the merge.
- Release-note writer. Changelog in, customer-facing notes out, plus a list of claims it could not verify (see the test below).
- Ops agent across tools. An internal assistant that updates CRM records, files helpdesk tickets, and drafts follow-ups, the kind of work AutomationBench measures.
- Research briefs. Long source packs condensed into a decision memo with the open questions listed.
What we saw in a small test
We gave openai/gpt-6-sol a three-line changelog and asked for customer-facing release notes, plus a list of claims it could not verify from the text.
It wrote three notes that stayed inside the source and did not invent a version number or a benefit. It also flagged two real gaps: the changelog does not say which SDK was bumped, and it gives no retry timing or limits.
That is one small task, not a benchmark. It does match the behavior OpenAI highlights: saying what it could not check.
When Should You Use GPT-6 Sol?
| Workload | Luna | Sol | Astra |
|---|---|---|---|
| Routine classification or extraction | Best fit | Usually more than needed | Not worth it |
| Everyday coding, reviews, single-repo changes | Possible for small edits | Best fit | For the hardest cases |
| Multi-step agent workflows across tools | Short sub-steps only | Best fit | When failure is costly |
| Long computer-use sessions | Limited | Workable | Strongest (72.6% on OSWorld 2.0) |
| Research and report writing | First drafts | Best fit | When depth decides the outcome |
The pattern: Sol is the default whenever the model has to decide something, and Luna takes the jobs where the answer is already constrained.
GPT-6 Sol vs Luna: When Is Sol Worth 20x?
Pay for Sol when a wrong step costs more than the tokens. A mislabeled ticket is cheap to fix. A wrong migration or a broken agent run is not.
On pure coding benchmarks the two are closer than the price suggests: 68.8% versus 66.6% on DeepSWE 1.1 at max effort. The difference shows up in positioning and judgment, not in that one score.
A common setup is Luna first, Sol on escalation. See the GPT-6 Luna guide for where Luna holds up on its own.
GPT-6 Sol vs Astra: When Is Astra Worth 5x?
Astra's lead is largest in computer use and long, open-ended work. On OSWorld 2.0, Astra scores 72.6% against Sol's 60.5%.
On AutomationBench, the gap runs the other way on cost: Sol at xhigh beats Astra at low effort, at about a quarter of the price per task.
So the question is not "which is smarter" but "which task fails on Sol." Run the task on Sol first, and move only the failures to Astra. The GPT-6 Astra guide covers what the flagship adds.
How to Access and Use GPT-6 Sol
Sol is available in ChatGPT Work and Codex on paid plans. In the API, you can call it directly on OpenAI as gpt-6-sol, or through an OpenAI-compatible gateway such as Modellix, where it is openai/gpt-6-sol. Two details matter either way:
- Tools need the Responses API. On Chat Completions, OpenAI supports function calling for Sol only when
reasoning_effortisnone. - Effort is a cost dial. The default is medium. Benchmark scores above use xhigh or max, which cost more per task.
The example below uses the OpenAI Python SDK with Modellix. Create a key in the Modellix console, then change two lines from a standard OpenAI setup: base_url and the model name. To call OpenAI directly instead, drop base_url and use gpt-6-sol.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://llm.modellix.ai/v1",
api_key=os.environ["MODELLIX_API_KEY"],
)
resp = client.responses.create(
model="openai/gpt-6-sol",
reasoning={"effort": "high"},
input="Review this diff and list any change that could break existing clients: ...",
)
print(resp.output_text)
Try GPT-6 Sol with One API Key for Every Model
Sol is rarely the only model in a production stack. Teams keep Luna for bulk work, Claude or Gemini for second opinions, and an image model for anything visual, each with its own key, bill, and SDK setup.
However, the Modellix key from the example above covers that whole stack. It also reaches GPT-6 Luna, GPT-6 Astra, Claude, Gemini, and Grok, so escalating a failed task to Astra is a one-string change to model.
Same tokens, lower bill
OpenAI's official API pricing lists GPT-6 Sol at $2 per million input tokens and $10 per million output tokens. For one call with 8,000 input tokens, 8,000 output tokens, and no cache, that is about $0.096 on OpenAI and $0.0864 on Modellix, as the Modellix LLM calculator shows below.
| Calls at 8,000 in + 8,000 out | OpenAI direct | Modellix | Saved |
|---|---|---|---|
| 1,000 | $96 | $86.40 | $9.60 |
| 10,000 | $960 | $864 | $96 |
| 100,000 | $9,600 | $8,640 | $960 |
| 1,000,000 | $96,000 | $86,400 | $9,600 |
At 1,000 calls that is $9.60 less, and at 1 million calls it is $9,600 less than OpenAI direct. This is a price comparison at the same billed token counts, not a quality benchmark, and it stays inside the 272K input tier.
One API for every model
Run GPT-6 Sol next to every other model on one key
GPT-6 Astra, Sol, and Luna, plus Claude, Gemini, Grok, DeepSeek, and Qwen. Switching models is a change to the model string.
OpenAI Chat Completions and Responses, plus Anthropic Messages. Works with the OpenAI and Anthropic SDKs, Codex, Claude Code, Cursor, and OpenCode.
GPT-6 and Claude models bill 10% under the vendor list price on the current rate card. Pay as you go, no subscription.
Every call is logged with tokens and cost, and can be tagged by end user, so you can see what an escalation actually cost.
Pin a fixed model ID for evaluations, or use a -latest alias that moves to the newest release.
The same account also calls image, video, and audio generation models through the Modellix media API.
The same key also runs image and video models. A release-notes workflow on Sol can hand the approved copy to GPT Image 2 for the announcement visual, or to Seedance 2.0 for a short product clip.
Pin openai/gpt-6-sol for evaluations, or use ~openai/gpt-sol-latest to pick up future Sol releases without a code change.
Try GPT-6 Sol
Create a Modellix key, run Sol on a task your team already knows, and escalate only the failures to Astra.
Get API Key



