
GPT-6 Luna is the smallest and cheapest model in OpenAI's GPT-6 family, released on September 22, 2026 alongside GPT-6 Sol. It costs $0.10 per million input tokens and $0.50 per million output tokens, twenty times less than Sol on both sides.
OpenAI describes it as its most efficient model "for focused, high-volume tasks." In practice that means classification, extraction, routing, and the short repeated calls inside an agent.
If the task needs several steps of real reasoning, Sol is usually the better buy. The rest of this page is about finding that line.
GPT-6 Luna at a Glance
| GPT-6 Luna | |
|---|---|
| Position | Volume tier of GPT-6, below Sol and Astra; replaces GPT-5.6 Luna |
| Best for | High-volume, well-defined tasks: labels, fields, routing, short summaries |
| Input price | $0.10 per 1M tokens ($0.01 cached) |
| Output price | $0.50 per 1M tokens |
| Context window | 1,050,000 tokens, up to 128,000 output tokens |
| Long-prompt rule | Above 272K input tokens, the whole request bills at 2x input and 1.5x output |
| Reasoning effort | none, low, medium (default), high, xhigh, max |
| Input types | Text and images in, text out |
| Tools | Function calling, structured outputs, web search, file search, computer use, MCP (Responses API) |
| Model ID | gpt-6-luna on OpenAI, openai/gpt-6-luna on Modellix |
| Step up | GPT-6 Sol for multi-step coding and agent work |
Specs and prices from OpenAI's model page.
GPT-6 Luna vs GPT-5.6 Luna
A lot of "GPT Luna" searches still point at the previous model, so here is the short version of what changed.
| GPT-5.6 Luna | GPT-6 Luna | |
|---|---|---|
| Input / output per 1M | $0.20 / $1.20 | $0.10 / $0.50 |
| AutomationBench (high effort) | Baseline | +5.4 points, 58% lower cost per task |
| Factuality | Baseline | At higher effort, matches GPT-5.6 Sol at about 1/100 of the cost |
The output price is the bigger change. Output dropped by 58%, and output tokens are usually what dominate a Luna bill when answers are longer than a label.
If you are on GPT-5.6 Sol only because Luna was not accurate enough, rerun those tasks on GPT-6 Luna first. That is the migration with the most room to save.
GPT-6 Luna Pricing
The list price is simple. The bill depends on three things: output length, cache hits, and whether a prompt crosses 272K tokens.
| Rate (per 1M tokens) | ≤272K input | >272K input |
|---|---|---|
| Input | $0.10 | $0.20 |
| Cached input read | $0.01 | $0.02 |
| Cache write | $0.125 | $0.25 |
| Output | $0.50 | $0.75 |
Batch and Flex run at half these rates on OpenAI. Fast mode doubles them.
What real workloads cost
We used the same token counts for all three GPT-6 tiers so the gap is easy to read.
| Workload | Luna | Sol | Astra |
|---|---|---|---|
| 10M input + 2M output a month | $2.00 | $40.00 | $200.00 |
| Same, 80% of input served from cache | $1.28 | $25.60 | $128.00 |
| One 400K-token prompt, 5K output | $0.084 | $1.68 | $8.38 |
These are our calculations from OpenAI list prices, excluding cache-write charges. The last row is billed at the >272K rates for every token, so it costs about twice what the same prompt would below the line.
A support queue of 100,000 tickets, each around 300 tokens in and 5 tokens out with reasoning off, costs about $3.25 on Luna. That is the kind of job the model exists for.
GPT-6 Luna Context Window
Luna reads up to 1,050,000 tokens and writes up to 128,000. That is the same window as Sol and Astra.
The number that matters for cost is 272K. Below it, you pay the base rate. One token above it, the whole request moves to 2x input and 1.5x output.
So the useful rule is to keep repeated calls under 272K and cache the shared prefix. Save the long prompts for the occasional whole-document pass, where Luna's low base price still keeps a 400K read under ten cents.
What Can You Build with GPT-6 Luna
Luna earns its place in products where the model makes thousands of small, checkable decisions a day.
- Support inbox triage. Every new ticket gets a label and a priority before a person opens it, and only the unclear ones reach a bigger model.
- Catalog and invoice extraction. An e-commerce back office turns supplier PDFs and emails into clean JSON fields with structured outputs.
- Moderation queues. A community app pre-sorts reports into "remove," "review," and "fine," so moderators start with the hard cases.
- Agent routers. Inside a coding or research assistant, Luna decides which tool or model handles the next step.
What we saw in a small test
We sent 12 hand-written support tickets to openai/gpt-6-luna through the Modellix gateway, each with a fixed label set: billing, bug, feature request, or other.
Luna labeled all 12 correctly, including two that could go either way ("My card was declined but the payment still shows as pending" and "Webhook retries stop after the first 500 error instead of backing off"). Each answer was five or six output tokens with nothing to strip.
Twelve tickets is a smoke test, not an accuracy rate. Before you route real traffic, run a few hundred tickets your team has already labeled.
When Luna is probably the wrong choice
- Multi-file code changes or long debugging sessions.
- Agent runs with many dependent steps, where one wrong decision compounds.
- Anything where a wrong answer is expensive and hard to spot.
For those, start with Sol. OpenAI positions Sol for "complex coding and agentic workflows" and Luna for focused tasks, and it has not published a head-to-head of the two on agent benchmarks.
GPT-6 Luna vs GPT-6 Sol
| Luna | Sol | |
|---|---|---|
| Price per 1M (in / out) | $0.10 / $0.50 | $2.00 / $10.00 |
| DeepSWE 1.1 (max effort, OpenAI) | 66.6% | 68.8% |
| Positioning | Focused, high-volume tasks | Complex coding and agentic workflows |
| Best fit | Labels, fields, routing, sub-steps | Multi-step work that needs judgment |
The coding numbers are close. The price is not.
That makes a two-tier setup the default for most teams: Luna handles the bulk, and only the uncertain or high-stakes cases escalate to Sol. The GPT-6 Sol guide covers where Sol earns its price.
How to Use GPT-6 Luna
You can call Luna directly on OpenAI's API as gpt-6-luna, or through an OpenAI-compatible gateway such as Modellix, where it is openai/gpt-6-luna. Either way, two details save debugging time:
- Use the Responses API for tools. OpenAI's Chat Completions supports function calling on Luna only with
reasoning_effortset tonone. - Set reasoning effort on purpose. The default is medium. For one-word labels,
lowornoneis usually enough and returns faster.
The example below uses the OpenAI Python SDK with Modellix. Create a key in the Modellix console, then change two lines from a standard OpenAI setup: base_url and the model name. To call OpenAI directly instead, drop base_url and use gpt-6-luna.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://llm.modellix.ai/v1",
api_key=os.environ["MODELLIX_API_KEY"],
)
resp = client.responses.create(
model="openai/gpt-6-luna",
reasoning={"effort": "low"},
input="Label this ticket as billing, bug, feature request, or other. "
"Reply with the label only. Ticket: I was charged twice in September.",
)
print(resp.output_text)
Try GPT-6 Luna with One API Key for Every Model
Most teams that adopt Luna do it to route cheap calls away from a bigger model. That usually means Luna on one account, Claude or Gemini on another, and a third integration for image generation.
However, the Modellix key from the example above is not a Luna-only key. It also reaches GPT-6 Sol, GPT-6 Astra, Claude, Gemini, and Grok, so moving a ticket from Luna to Sol is a one-string change to model.
Same tokens, lower bill
OpenAI's official API pricing lists GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output tokens. For one call with 8,000 input tokens, 8,000 output tokens, and no cache, that is about $0.0048 on OpenAI and $0.00432 on Modellix, as the Modellix LLM calculator shows below.
| Calls at 8,000 in + 8,000 out | OpenAI direct | Modellix | Saved |
|---|---|---|---|
| 1,000 | $4.80 | $4.32 | $0.48 |
| 10,000 | $48 | $43.20 | $4.80 |
| 100,000 | $480 | $432 | $48 |
| 1,000,000 | $4,800 | $4,320 | $480 |
At 1 million calls with that mix, Modellix bills $480 less than OpenAI direct. This is a price comparison at the same billed token counts, not a quality benchmark, and it stays inside the 272K input tier.
One API for every model
Run GPT-6 Luna next to every other model on one key
GPT-6 Astra, Sol, and Luna, plus Claude, Gemini, Grok, DeepSeek, and Qwen. Switching models is a change to the model string.
OpenAI Chat Completions and Responses, plus Anthropic Messages. Works with the OpenAI and Anthropic SDKs, Codex, Claude Code, Cursor, and OpenCode.
GPT-6 and Claude models bill 10% under the vendor list price on the current rate card. Pay as you go, no subscription.
Every call is logged with tokens and cost, and can be tagged by end user, so you can see what an escalation actually cost.
Pin a fixed model ID for evaluations, or use a -latest alias that moves to the newest release.
The same account also calls image, video, and audio generation models through the Modellix media API.
The same key reaches image and video models. A support team that uses Luna to spot recurring ticket themes can turn them into help-center visuals with GPT Image 2, without another account.
Pin openai/gpt-6-luna for evaluations, or use ~openai/gpt-luna-latest to pick up future Luna releases without a code change.
Try GPT-6 Luna
Create a Modellix key, run Luna on a batch you have already labeled, and move only the hard cases to Sol.
Get API Key



