Guides

GPT-6 Sol: Pricing, Benchmarks & When to Use It

GPT-6 Sol costs $2/$10 per 1M tokens, half of GPT-5.6 Sol. See what its benchmarks measure, real workload costs, and when to pick Sol over Luna or Astra.

37 min read
OpenAIGPT-6LLM API
Modellix Team
Written byModellix TeamOfficial
GPT-6 Sol: Pricing, Benchmarks & When to Use It

GPT-6 Sol is the middle model in OpenAI's GPT-6 family, released on September 22, 2026. It sits between Luna, the cheap volume tier, and Astra, the flagship.

It costs $2 per million input tokens and $10 per million output tokens. That is half the price of GPT-5.6 Sol and one fifth of Astra.

OpenAI built it "to power complex coding and agentic workflows." The short answer for most teams: use Sol as the default for work that needs judgment, and treat Astra as the exception you have to justify.

GPT-6 Sol at a Glance

GPT-6 Sol
Position Middle GPT-6 tier, above GPT-6 Luna and below GPT-6 Astra; replaces GPT-5.6 Sol
Best for Coding agents, multi-step business workflows, research and synthesis
Input price $2.00 per 1M tokens ($0.20 cached)
Output price $10.00 per 1M tokens
Context window 1,050,000 tokens, up to 128,000 output tokens
Long-prompt rule Above 272K input tokens, the whole request bills at 2x input and 1.5x output
Reasoning effort none, low, medium (default), high, xhigh, max
Tools Function calling, structured outputs, web search, computer use, hosted shell, apply patch, MCP (Responses API)
Model ID gpt-6-sol on OpenAI, openai/gpt-6-sol on Modellix

Specs and prices from OpenAI's model page.

GPT-6 Sol vs GPT-5.6 Sol

GPT-5.6 Sol GPT-6 Sol
Input / output per 1M $4 / $20 $2 / $10
FrontierCode 1.1 Main Baseline "Improves substantially," matching Fable 5.1 (xhigh)
Factuality (OpenAI internal) Baseline About half as many mistakes
Answer style Longer, restates details Shorter, says what it checked

The price cut applies to every token, so an unchanged workload costs half as much.

One number does not move in Sol's favor. OpenAI's Astra launch table lists GPT-5.6 Sol at 72.7% on DeepSWE 1.1, while the Sol post reports GPT-6 Sol at 68.8% at max effort. The two posts may not use the same setup, but it is a reason to test coding-heavy workloads before switching blind.

If you run GPT-5.6 Sol today, switching is mostly a model-ID change. Rerun your own evaluation anyway, because shorter answers can break parsers that expect the old format.

GPT-6 Sol Pricing

Rate (per 1M tokens) ≤272K input >272K input
Input $2.00 $4.00
Cached input read $0.20 $0.40
Cache write $2.50 $5.00
Output $10.00 $15.00

Batch and Flex are half price on OpenAI. Fast mode is double.

The same workload on all three GPT-6 models

Workload Luna Sol Astra
10M input + 2M output a month $2.00 $40.00 $200.00
Same, 80% of input served from cache $1.28 $25.60 $128.00
One 400K-token prompt, 5K output $0.084 $1.68 $8.38

These are our calculations from OpenAI list prices, excluding cache-write charges.

Caching matters more on Sol than on Luna in absolute terms. Agents that resend the same system prompt and tool definitions every turn can cut a third off the bill at an 80% hit rate, and OpenAI says GPT-6 caching now holds across changes in reasoning effort and tool availability.

For Astra's full rate card, including cache and long-context tiers, see our GPT-6 Astra pricing breakdown.

GPT-6 Sol Benchmarks

All numbers below are from OpenAI's launch post.

Benchmark What it measures GPT-6 Sol Compared with Result Cost per task
AutomationBench 1.0.6 End-to-end business workflows across 47 tools 33.2% xhigh GPT-6 Astra (low) 30.3%, Claude Opus 5 (max) 26.9% Beats both $0.27; Astra costs 3.9x, Opus 5 11.1x
Agents' Last Exam Long professional tasks across 55 sub-industries 56.4% max Claude Opus 5's best score Beats Opus 5 60% lower than Opus 5
OSWorld 2.0 offline Computer-use workflows, partial reward 60.5% xhigh Claude Opus 5 (medium) 60.3% Ties About 80% lower
FrontierCode 1.1 Main Whether code is ready to merge, not just correct Not published Claude Fable 5.1 (xhigh) Matches "Much lower" (not quantified)
DeepSWE 1.1 Long software engineering tasks in real codebases 68.8% max Claude Fable 5 best, 69.9% (xhigh) 1.1 pts behind About 80% lower

Two things to keep in mind when reading this table.

First, every Sol score is at xhigh or max effort. Those settings produce more reasoning tokens, and reasoning tokens bill as output.

Second, the comparisons are cost-per-task claims as much as score claims. Sol's case is "close to the frontier for much less," not "best in class."

What Can You Build with GPT-6 Sol

Sol fits products where the model has to reason over several steps and the output goes to a customer or a codebase.

  • Pull request reviewer. A bot that reads a diff, flags breaking changes, and explains why, before a human approves the merge.
  • Release-note writer. Changelog in, customer-facing notes out, plus a list of claims it could not verify (see the test below).
  • Ops agent across tools. An internal assistant that updates CRM records, files helpdesk tickets, and drafts follow-ups, the kind of work AutomationBench measures.
  • Research briefs. Long source packs condensed into a decision memo with the open questions listed.

What we saw in a small test

We gave openai/gpt-6-sol a three-line changelog and asked for customer-facing release notes, plus a list of claims it could not verify from the text.

It wrote three notes that stayed inside the source and did not invent a version number or a benefit. It also flagged two real gaps: the changelog does not say which SDK was bumped, and it gives no retry timing or limits.

That is one small task, not a benchmark. It does match the behavior OpenAI highlights: saying what it could not check.

When Should You Use GPT-6 Sol?

Workload Luna Sol Astra
Routine classification or extraction Best fit Usually more than needed Not worth it
Everyday coding, reviews, single-repo changes Possible for small edits Best fit For the hardest cases
Multi-step agent workflows across tools Short sub-steps only Best fit When failure is costly
Long computer-use sessions Limited Workable Strongest (72.6% on OSWorld 2.0)
Research and report writing First drafts Best fit When depth decides the outcome

The pattern: Sol is the default whenever the model has to decide something, and Luna takes the jobs where the answer is already constrained.

GPT-6 Sol vs Luna: When Is Sol Worth 20x?

Pay for Sol when a wrong step costs more than the tokens. A mislabeled ticket is cheap to fix. A wrong migration or a broken agent run is not.

On pure coding benchmarks the two are closer than the price suggests: 68.8% versus 66.6% on DeepSWE 1.1 at max effort. The difference shows up in positioning and judgment, not in that one score.

A common setup is Luna first, Sol on escalation. See the GPT-6 Luna guide for where Luna holds up on its own.

GPT-6 Sol vs Astra: When Is Astra Worth 5x?

Astra's lead is largest in computer use and long, open-ended work. On OSWorld 2.0, Astra scores 72.6% against Sol's 60.5%.

On AutomationBench, the gap runs the other way on cost: Sol at xhigh beats Astra at low effort, at about a quarter of the price per task.

So the question is not "which is smarter" but "which task fails on Sol." Run the task on Sol first, and move only the failures to Astra. The GPT-6 Astra guide covers what the flagship adds.

How to Access and Use GPT-6 Sol

Sol is available in ChatGPT Work and Codex on paid plans. In the API, you can call it directly on OpenAI as gpt-6-sol, or through an OpenAI-compatible gateway such as Modellix, where it is openai/gpt-6-sol. Two details matter either way:

  1. Tools need the Responses API. On Chat Completions, OpenAI supports function calling for Sol only when reasoning_effort is none.
  2. Effort is a cost dial. The default is medium. Benchmark scores above use xhigh or max, which cost more per task.

The example below uses the OpenAI Python SDK with Modellix. Create a key in the Modellix console, then change two lines from a standard OpenAI setup: base_url and the model name. To call OpenAI directly instead, drop base_url and use gpt-6-sol.

python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://llm.modellix.ai/v1",
    api_key=os.environ["MODELLIX_API_KEY"],
)

resp = client.responses.create(
    model="openai/gpt-6-sol",
    reasoning={"effort": "high"},
    input="Review this diff and list any change that could break existing clients: ...",
)
print(resp.output_text)

Try GPT-6 Sol with One API Key for Every Model

Sol is rarely the only model in a production stack. Teams keep Luna for bulk work, Claude or Gemini for second opinions, and an image model for anything visual, each with its own key, bill, and SDK setup.

However, the Modellix key from the example above covers that whole stack. It also reaches GPT-6 Luna, GPT-6 Astra, Claude, Gemini, and Grok, so escalating a failed task to Astra is a one-string change to model.

Same tokens, lower bill

OpenAI's official API pricing lists GPT-6 Sol at $2 per million input tokens and $10 per million output tokens. For one call with 8,000 input tokens, 8,000 output tokens, and no cache, that is about $0.096 on OpenAI and $0.0864 on Modellix, as the Modellix LLM calculator shows below.

gpt 6 sol price calculator
Modellix LLM calculator for 8,000 input and 8,000 output tokens. Enter your own token mix on the live page.
Calls at 8,000 in + 8,000 out OpenAI direct Modellix Saved
1,000 $96 $86.40 $9.60
10,000 $960 $864 $96
100,000 $9,600 $8,640 $960
1,000,000 $96,000 $86,400 $9,600

At 1,000 calls that is $9.60 less, and at 1 million calls it is $9,600 less than OpenAI direct. This is a price comparison at the same billed token counts, not a quality benchmark, and it stays inside the 272K input tier.

One API for every model

Run GPT-6 Sol next to every other model on one key

One key, every model

GPT-6 Astra, Sol, and Luna, plus Claude, Gemini, Grok, DeepSeek, and Qwen. Switching models is a change to the model string.

Drop-in for your client

OpenAI Chat Completions and Responses, plus Anthropic Messages. Works with the OpenAI and Anthropic SDKs, Codex, Claude Code, Cursor, and OpenCode.

Below list price

GPT-6 and Claude models bill 10% under the vendor list price on the current rate card. Pay as you go, no subscription.

Cost per request, in the logs

Every call is logged with tokens and cost, and can be tagged by end user, so you can see what an escalation actually cost.

Pin or follow

Pin a fixed model ID for evaluations, or use a -latest alias that moves to the newest release.

Media on the same key

The same account also calls image, video, and audio generation models through the Modellix media API.

The same key also runs image and video models. A release-notes workflow on Sol can hand the approved copy to GPT Image 2 for the announcement visual, or to Seedance 2.0 for a short product clip.

Pin openai/gpt-6-sol for evaluations, or use ~openai/gpt-sol-latest to pick up future Sol releases without a code change.