Guides

GPT-6 Luna: Pricing, Context Window & How to Use It

GPT-6 Luna costs $0.10 per 1M input tokens with a 1.05M context window. See workload costs, Luna vs Sol, what changed from GPT-5.6, and how to call it.

29 min read
OpenAIGPT-6LLM API
Modellix Team
Written byModellix TeamOfficial
GPT-6 Luna: Pricing, Context Window & How to Use It

GPT-6 Luna is the smallest and cheapest model in OpenAI's GPT-6 family, released on September 22, 2026 alongside GPT-6 Sol. It costs $0.10 per million input tokens and $0.50 per million output tokens, twenty times less than Sol on both sides.

OpenAI describes it as its most efficient model "for focused, high-volume tasks." In practice that means classification, extraction, routing, and the short repeated calls inside an agent.

If the task needs several steps of real reasoning, Sol is usually the better buy. The rest of this page is about finding that line.

GPT-6 Luna at a Glance

GPT-6 Luna
Position Volume tier of GPT-6, below Sol and Astra; replaces GPT-5.6 Luna
Best for High-volume, well-defined tasks: labels, fields, routing, short summaries
Input price $0.10 per 1M tokens ($0.01 cached)
Output price $0.50 per 1M tokens
Context window 1,050,000 tokens, up to 128,000 output tokens
Long-prompt rule Above 272K input tokens, the whole request bills at 2x input and 1.5x output
Reasoning effort none, low, medium (default), high, xhigh, max
Input types Text and images in, text out
Tools Function calling, structured outputs, web search, file search, computer use, MCP (Responses API)
Model ID gpt-6-luna on OpenAI, openai/gpt-6-luna on Modellix
Step up GPT-6 Sol for multi-step coding and agent work

Specs and prices from OpenAI's model page.

GPT-6 Luna vs GPT-5.6 Luna

A lot of "GPT Luna" searches still point at the previous model, so here is the short version of what changed.

GPT-5.6 Luna GPT-6 Luna
Input / output per 1M $0.20 / $1.20 $0.10 / $0.50
AutomationBench (high effort) Baseline +5.4 points, 58% lower cost per task
Factuality Baseline At higher effort, matches GPT-5.6 Sol at about 1/100 of the cost

The output price is the bigger change. Output dropped by 58%, and output tokens are usually what dominate a Luna bill when answers are longer than a label.

If you are on GPT-5.6 Sol only because Luna was not accurate enough, rerun those tasks on GPT-6 Luna first. That is the migration with the most room to save.

GPT-6 Luna Pricing

The list price is simple. The bill depends on three things: output length, cache hits, and whether a prompt crosses 272K tokens.

Rate (per 1M tokens) ≤272K input >272K input
Input $0.10 $0.20
Cached input read $0.01 $0.02
Cache write $0.125 $0.25
Output $0.50 $0.75

Batch and Flex run at half these rates on OpenAI. Fast mode doubles them.

What real workloads cost

We used the same token counts for all three GPT-6 tiers so the gap is easy to read.

Workload Luna Sol Astra
10M input + 2M output a month $2.00 $40.00 $200.00
Same, 80% of input served from cache $1.28 $25.60 $128.00
One 400K-token prompt, 5K output $0.084 $1.68 $8.38

These are our calculations from OpenAI list prices, excluding cache-write charges. The last row is billed at the >272K rates for every token, so it costs about twice what the same prompt would below the line.

A support queue of 100,000 tickets, each around 300 tokens in and 5 tokens out with reasoning off, costs about $3.25 on Luna. That is the kind of job the model exists for.

GPT-6 Luna Context Window

Luna reads up to 1,050,000 tokens and writes up to 128,000. That is the same window as Sol and Astra.

The number that matters for cost is 272K. Below it, you pay the base rate. One token above it, the whole request moves to 2x input and 1.5x output.

So the useful rule is to keep repeated calls under 272K and cache the shared prefix. Save the long prompts for the occasional whole-document pass, where Luna's low base price still keeps a 400K read under ten cents.

What Can You Build with GPT-6 Luna

Luna earns its place in products where the model makes thousands of small, checkable decisions a day.

  • Support inbox triage. Every new ticket gets a label and a priority before a person opens it, and only the unclear ones reach a bigger model.
  • Catalog and invoice extraction. An e-commerce back office turns supplier PDFs and emails into clean JSON fields with structured outputs.
  • Moderation queues. A community app pre-sorts reports into "remove," "review," and "fine," so moderators start with the hard cases.
  • Agent routers. Inside a coding or research assistant, Luna decides which tool or model handles the next step.

What we saw in a small test

We sent 12 hand-written support tickets to openai/gpt-6-luna through the Modellix gateway, each with a fixed label set: billing, bug, feature request, or other.

Luna labeled all 12 correctly, including two that could go either way ("My card was declined but the payment still shows as pending" and "Webhook retries stop after the first 500 error instead of backing off"). Each answer was five or six output tokens with nothing to strip.

Twelve tickets is a smoke test, not an accuracy rate. Before you route real traffic, run a few hundred tickets your team has already labeled.

When Luna is probably the wrong choice

  • Multi-file code changes or long debugging sessions.
  • Agent runs with many dependent steps, where one wrong decision compounds.
  • Anything where a wrong answer is expensive and hard to spot.

For those, start with Sol. OpenAI positions Sol for "complex coding and agentic workflows" and Luna for focused tasks, and it has not published a head-to-head of the two on agent benchmarks.

GPT-6 Luna vs GPT-6 Sol

Luna Sol
Price per 1M (in / out) $0.10 / $0.50 $2.00 / $10.00
DeepSWE 1.1 (max effort, OpenAI) 66.6% 68.8%
Positioning Focused, high-volume tasks Complex coding and agentic workflows
Best fit Labels, fields, routing, sub-steps Multi-step work that needs judgment

The coding numbers are close. The price is not.

That makes a two-tier setup the default for most teams: Luna handles the bulk, and only the uncertain or high-stakes cases escalate to Sol. The GPT-6 Sol guide covers where Sol earns its price.

How to Use GPT-6 Luna

You can call Luna directly on OpenAI's API as gpt-6-luna, or through an OpenAI-compatible gateway such as Modellix, where it is openai/gpt-6-luna. Either way, two details save debugging time:

  1. Use the Responses API for tools. OpenAI's Chat Completions supports function calling on Luna only with reasoning_effort set to none.
  2. Set reasoning effort on purpose. The default is medium. For one-word labels, low or none is usually enough and returns faster.

The example below uses the OpenAI Python SDK with Modellix. Create a key in the Modellix console, then change two lines from a standard OpenAI setup: base_url and the model name. To call OpenAI directly instead, drop base_url and use gpt-6-luna.

python
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://llm.modellix.ai/v1",
    api_key=os.environ["MODELLIX_API_KEY"],
)

resp = client.responses.create(
    model="openai/gpt-6-luna",
    reasoning={"effort": "low"},
    input="Label this ticket as billing, bug, feature request, or other. "
          "Reply with the label only. Ticket: I was charged twice in September.",
)
print(resp.output_text)

Try GPT-6 Luna with One API Key for Every Model

Most teams that adopt Luna do it to route cheap calls away from a bigger model. That usually means Luna on one account, Claude or Gemini on another, and a third integration for image generation.

However, the Modellix key from the example above is not a Luna-only key. It also reaches GPT-6 Sol, GPT-6 Astra, Claude, Gemini, and Grok, so moving a ticket from Luna to Sol is a one-string change to model.

Same tokens, lower bill

OpenAI's official API pricing lists GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output tokens. For one call with 8,000 input tokens, 8,000 output tokens, and no cache, that is about $0.0048 on OpenAI and $0.00432 on Modellix, as the Modellix LLM calculator shows below.

gpt 6 luna price calculator
Modellix LLM calculator for 8,000 input and 8,000 output tokens. Enter your own token mix on the live page.
Calls at 8,000 in + 8,000 out OpenAI direct Modellix Saved
1,000 $4.80 $4.32 $0.48
10,000 $48 $43.20 $4.80
100,000 $480 $432 $48
1,000,000 $4,800 $4,320 $480

At 1 million calls with that mix, Modellix bills $480 less than OpenAI direct. This is a price comparison at the same billed token counts, not a quality benchmark, and it stays inside the 272K input tier.

One API for every model

Run GPT-6 Luna next to every other model on one key

One key, every model

GPT-6 Astra, Sol, and Luna, plus Claude, Gemini, Grok, DeepSeek, and Qwen. Switching models is a change to the model string.

Drop-in for your client

OpenAI Chat Completions and Responses, plus Anthropic Messages. Works with the OpenAI and Anthropic SDKs, Codex, Claude Code, Cursor, and OpenCode.

Below list price

GPT-6 and Claude models bill 10% under the vendor list price on the current rate card. Pay as you go, no subscription.

Cost per request, in the logs

Every call is logged with tokens and cost, and can be tagged by end user, so you can see what an escalation actually cost.

Pin or follow

Pin a fixed model ID for evaluations, or use a -latest alias that moves to the newest release.

Media on the same key

The same account also calls image, video, and audio generation models through the Modellix media API.

The same key reaches image and video models. A support team that uses Luna to spot recurring ticket themes can turn them into help-center visuals with GPT Image 2, without another account.

Pin openai/gpt-6-luna for evaluations, or use ~openai/gpt-luna-latest to pick up future Luna releases without a code change.