If you searched for the best LLMs, you probably wanted a name. The honest answer is that the name depends on three numbers you can look up in about ten minutes, and on nothing else: how much text the model can accept at once, what kinds of input it takes besides text, and what it charges per million tokens — including whether that price changes once your input crosses a threshold.
That is the whole decision. It is not a shortlist problem, it is a filter problem. Apply the filters in the right order and a catalogue of thirty models collapses to two or three candidates without anybody’s opinion entering the loop.
We run Modellix, which is the API gateway whose catalogue this page compares, so we have a commercial interest in you reading it — that is stated up front rather than buried in a footer, and everything below is checkable against pages you can open yourself. The catalogue figures were captured on September 17, 2026 from the live model table, and the commands at the end of this page let you re-pull every number instead of trusting ours.
There is no best LLM in 2026, and the question is built on a broken unit
“Best” needs a unit. For a chat subscription the unit is an app you open and a monthly price. For an API the unit is a model you call, priced per million tokens, and constrained by a context window, an input modality list, and a pricing mode. Those are different units, which is why the same question produces answers that cannot be compared to each other.
Here is the shape of the correct answer, before any detail: no single best LLM exists, because “best” without a workload is undecidable, and the three attributes that do decide it are published for every model in the catalogue.
Watch how fast the filter works. If your input includes video, you are choosing among 8 of the 30 models. If it includes audio, among 4. If you need a 1.3-million-token window, you are looking at one model. If your workload re-sends a large stable prefix, the cached-input rate matters more than the headline input rate, and that narrows the field again. None of those filters requires a benchmark, a preference vote, or someone’s hands-on opinion — and none of them is visible in a ranked list of model names.
This is also why a page cannot honestly hand you a single winner. Any name we printed would be true only for one workload, and wrong for the next reader. What we can do is hand you the filter and the table it runs against.
Why every “best LLMs” list disagrees — it is an intent split, not a data dispute
Read the pages that rank first for this term and you will notice they are not arguing with each other. They are answering different questions, for different readers, and the two groups almost never overlap.
One group is written for people choosing a chat product. Their unit of comparison is the app: ChatGPT, Claude, Gemini, and whichever subscription wraps them. Their dimensions are conversational quality, personal-assistant usefulness, and how good the free tier is. The other group is written for people choosing a model to call — price per million tokens, tokens per second, time to first token, context window, and which models other developers are actually routing traffic to. Leaderboard and usage-data pages sit in this second group.
Both are legitimate. The confusion is that they share a keyword. If you arrive at a page that tells you GPT-4o is a strong all-round choice, you are reading a document written against a 2024–2025 model generation, and it cannot answer a 2026 integration question — several of the 2026-titled lists we read still recommend that generation, which is the single easiest gap to notice. Ranked tables also have a subtler limit: they tell you who is first, and stop there. They do not tell you what a month of that choice costs on your traffic, whether the model takes an image input, or what happens to the bill when a request grows past a tier boundary. Those are the questions that come immediately after the ranking, and they are the ones this page is built around.
So if you want the current frontier ranking, a live leaderboard is the right tool and we will not pretend to replace it — this page will not republish scores we did not measure. If you want the decision that follows the ranking, keep reading.
The 30-model table: price, context and modality in one place
Below is the live model table as it stood on September 17, 2026: 30 language models under 9 provider labels (Anthropic, DeepSeek, Google, Moonshot, OpenAI, Qwen, xAI, ZAI, and Modellix’s own modellix-ai entry). Prices are USD per 1M tokens, exactly as the vendor’s table lists them; where two numbers appear as $X → $Y, the first is the upstream list price and the second is the rate the gateway bills.
Five columns, sorted by provider. Everything here is re-derivable from the model table and the commands in the last section.
| Model ID | Context (tokens) | Input modalities | Input $/1M | Output $/1M |
|---|---|---|---|---|
anthropic/claude-fable-5.1 |
1,000,000 | text / image | $10.00 → $9.00 | $50.00 → $45.00 |
anthropic/claude-haiku-4.5 |
200,000 | text / image | $1.00 → $0.90 | $5.00 → $4.50 |
anthropic/claude-opus-5 |
1,000,000 | text / image | $5.00 → $4.50 | $25.00 → $22.50 |
anthropic/claude-sonnet-5 |
1,000,000 | text / image | $2.00 → $1.80 | $10.00 → $9.00 |
deepseek/deepseek-v4-flash |
1,050,000 | text | $0.44 → $0.40 | $1.32 → $1.19 |
deepseek/deepseek-v4-flash-vision |
1,050,000 | text / image | $0.44 | $1.32 |
deepseek/deepseek-v4-pro |
1,050,000 | text | $1.33 → $1.20 | $3.96 → $3.56 |
deepseek/deepseek-v4.1-flash |
1,000,000 | text / image | $0.30 | $1.20 |
google/gemini-3.1-pro |
1,050,000 | text / image / video / audio | $2.00 → $1.80 (≤200,000) $4.00 → $3.60 (over 200,000) |
$12.00 → $10.80 (≤200,000) $18.00 → $16.20 (over 200,000) |
google/gemini-3.6-flash |
1,050,000 | text / image / video / audio | $1.50 → $1.35 | $7.50 → $6.75 |
google/gemini-3.7-flash |
1,050,000 | text / image / video / audio | $1.50 → $0.75 | $7.50 → $3.75 |
google/gemini-3.8-flash |
1,000,000 | text / image / video / audio | $1.50 → $0.75 | $7.50 → $3.75 |
modellix-ai/free-llm |
200,000 | text / image / video | $0.00 | $0.00 |
moonshot/kimi-k2.7 |
262,000 | text / image | $0.95 | $4.00 |
moonshot/kimi-k3 |
1,050,000 | text / image | $3.00 | $15.00 |
openai/gpt-5.5 |
1,050,000 | text / image | $5.00 → $4.50 (≤272,000) $10.00 → $9.00 (over 272,000) |
$30.00 → $27.00 (≤272,000) $45.00 → $40.50 (over 272,000) |
openai/gpt-5.6-luna |
1,050,000 | text / image | $0.20 → $0.18 (≤272,000) $0.40 → $0.36 (over 272,000) |
$1.20 → $1.08 (≤272,000) $1.80 → $1.62 (over 272,000) |
openai/gpt-5.6-sol |
1,050,000 | text / image | $4.00 → $3.60 (≤272,000) $8.00 → $7.20 (over 272,000) |
$20.00 → $18.00 (≤272,000) $30.00 → $27.00 (over 272,000) |
openai/gpt-5.6-terra |
1,050,000 | text / image | $2.00 → $1.80 (≤272,000) $4.00 → $3.60 (over 272,000) |
$12.00 → $10.80 (≤272,000) $18.00 → $16.20 (over 272,000) |
openai/gpt-6-astra |
1,050,000 | text / image | $10.00 → $9.00 (≤272,000) $20.00 → $18.00 (over 272,000) |
$50.00 → $45.00 (≤272,000) $75.00 → $67.50 (over 272,000) |
qwen/qwen3.7-max |
1,000,000 | text | $1.48 | $4.42 |
qwen/qwen3.7-plus |
1,000,000 | text / image | $0.32 | $1.28 |
qwen/qwen3.8-flash |
1,000,000 | text / image / video | $0.16 | $0.47 |
qwen/qwen3.8-max |
1,000,000 | text / image / video | $2.00 | $6.00 |
xai/grok-4.5 |
500,000 | text / image | $2.00 (≤200,000) $4.00 (over 200,000) |
$6.00 (≤200,000) $12.00 (over 200,000) |
xai/grok-4.6 |
500,000 | text / image | $2.00 (≤200,000) $4.00 (over 200,000) |
$6.00 (≤200,000) $12.00 (over 200,000) |
zai/glm-4.7-flash |
203,000 | text | $0.00 | $0.00 |
zai/glm-5.2 |
1,050,000 | text | $1.00 → $0.90 | $4.40 → $3.96 |
zai/glm-5.3 |
1,050,000 | text | $1.00 → $0.90 | $4.40 → $3.96 |
zai/glm-5.3-flash |
1,310,720 | text / image / video | $0.15 → $0.14 | $0.50 → $0.45 |
Two things in that table deserve to be said plainly before anyone reads a “best” ranking into it. The rows are expanded model by model in the 28-model pricing comparison behind this table, which covers the same rate columns in more depth.
First, the spread is three orders of magnitude wide. The cheapest paid entry is zai/glm-5.3-flash at $0.14 per 1M input and $0.45 per 1M output; the most expensive first-tier rate is $9.00 input and $45.00 output for anthropic/claude-fable-5.1, matched by openai/gpt-6-astra. On the same workload, that is the difference between cents and tens of dollars — which is why “which is better” is less useful than “which fits”.
Second, we are not going to claim the gateway is the cheapest route to any of them. The first number in each two-price cell — shown struck through on the live table — is the upstream list price as the vendor publishes it; where the right-hand number is lower, that is the rate this billing path applies. Where the two match, there is no difference at all — 12 of the 30 rows show a single price. A two-price cell is not a claim of being cheapest; it is a statement of what your invoice will show. And as the next sections show, one of the largest percentage gaps in the table has nothing to do with us.
Watch the output column, not the input column. Output tokens are where the multiples live. Across the paid entries, the first-tier input rate runs from $0.14 to $9.00 per million — roughly a 64× spread — while the output rate runs from $0.45 to $45.00, a 100× spread. Count the second tiers as well and it widens further: $18.00 input and $67.50 output, about 130× and 150× against the cheapest entry. A model with a cheap input rate and an expensive output rate is a bad fit for anything that generates long answers.
Filter one: input modality eliminates the most models fastest
Start here because it is the only filter that is a hard yes/no. A model either accepts an image in the request or it does not, and there is no partial credit.
Across the 30 models: 24 accept image input, 8 accept video, 4 accept audio, and 6 are text-only. The counts nest rather than partition — a model in the video group also accepts images. Audio input is the narrowest: all four are Google Gemini models (gemini-3.1-pro, gemini-3.6-flash, gemini-3.7-flash, gemini-3.8-flash). The six text-only entries are deepseek/deepseek-v4-flash, deepseek/deepseek-v4-pro, qwen/qwen3.7-max, zai/glm-4.7-flash, zai/glm-5.2, and zai/glm-5.3.
That single line rules out one fifth of the catalogue for a document-processing pipeline, and 26 of the 30 models for one that transcribes or reasons over audio.
Two details matter once you move from the table to a request. Field names follow the protocol you call, not a Modellix-specific schema, and they are not interchangeable between Chat Completions, Responses, and Messages — an image lives in image_url.url on the Chat Completions path, in input_image.image_url on the Responses path, and in image.source.url on the Messages path. The Media File API’s returned URL is the recommended way to hand over an upload, and local filesystem paths are ignored. If you send media and want to confirm it was actually consumed, check the usage block in the response rather than assuming.
The protocol split behind the modality counts: the same image input uses a different field on each of the three endpoints, and the API reference states that the Messages path has no video or audio block. Table captured from the gateway’s API reference on September 17, 2026.
One bound worth stating because it is easy to misread: every one of these 30 models returns text only. Inputs are multimodal; outputs are not. This is a text gateway — the image and video generation models in the full model catalogue beyond these 30 live on a different host with a different request shape, and mixing the two hosts or their fields is the most common integration mistake the documentation warns about.
Filter two: context window, and the priced tier most tables hide
Context window is the second filter, and it is the one most readers get half right. A window of 1,050,000 tokens says how much the model will accept. It says nothing about what it charges for accepting it — that is a separate mechanism, and it is where budgets break.
The window range across the catalogue runs from 200,000 tokens (anthropic/claude-haiku-4.5 and modellix-ai/free-llm) to 1,310,720 tokens (zai/glm-5.3-flash). Fourteen entries sit at 1,050,000, nine at exactly a million, two at 500,000, one at 1,310,720, and the remainder below that.
Now the part the tables hide. Eight of the 30 models do not have one price — they have two or more input-context tiers, and the tier is selected automatically by how many input tokens that specific request carries. The thresholds are not arbitrary and they are not the context window:
| Tiered models | Tier boundary (input tokens) | Tier 1 input → output | Tier 2 input → output |
|---|---|---|---|
google/gemini-3.1-pro, xai/grok-4.5, xai/grok-4.6 |
200,000 | $1.80 → $10.80 / $2.00 → $6.00 | $3.60 → $16.20 / $4.00 → $12.00 |
openai/gpt-5.5, gpt-5.6-sol, gpt-6-astra |
272,000 | $4.50 → $27.00 / $3.60 → $18.00 / $9.00 → $45.00 | $9.00 → $40.50 / $7.20 → $27.00 / $18.00 → $67.50 |
openai/gpt-5.6-luna, gpt-5.6-terra |
272,000 | $0.18 → $1.08 / $1.80 → $10.80 | $0.36 → $1.62 / $3.60 → $16.20 |
The Gemini 3.1 Pro row and the xAI rows carry a second combined cell for readability; the pattern is consistent — cross the boundary and the input rate doubles; on the six non-xAI tiered models the output rate rises by half as well.
The practical consequence is that the tier boundary is a design constraint, not a footnote. A request carrying 300,000 input tokens into openai/gpt-5.6-sol does not bill at the $3.60 rate; it bills at $7.20 input and $27.00 output, because the tier follows the request. If your pipeline concatenates retrieved documents into one call, the call length is a pricing decision. Splitting a 300,000-token request into two 150,000-token requests keeps both inside the first tier and is often cheaper than one call — but it doubles request count, so check it against your own numbers rather than taking our word for it.
The step, not the slope: a flat-rate model’s per-request cost rises linearly with size, while a tiered model jumps when the request crosses its boundary — at 200,000 input tokens for Gemini 3.1 Pro and Grok, at 272,000 for the OpenAI tiers. Values are computed from the published rates captured September 17, 2026, with no cached tokens assumed.
Twenty-two entries are flat-rate: one input price, one output price, no boundary to cross. For an unpredictable workload, a slightly higher flat rate is often the more boring and more budgetable choice — which is exactly the trade a ranked list cannot express, and the whole subject of the 272K tier that reprices whole requests.
Note the direction of the risk. Flat-rate models cannot surprise you on a long request. Tiered models can, and the surprise is a doubling of the input rate, not a percentage.
The chart below puts both axes together: input rate against context window, with output rate as the bubble size. The cheap-and-large corner is where most retrieval pipelines want to be, and the entries there are not the ones a quality ranking would put first.
Filter three: price mechanics — flat rate, tiers, cached reads, and one 50% label that is not ours
You now have a candidate set. This filter decides which of those candidates is affordable on your traffic, and it requires reading the price table the way the billing engine does rather than the way a summary does.
In the live table, four rate columns matter, not two: input, output, cached read, and cached write. The cached columns are the reason a workload with a large stable prefix behaves nothing like a workload with fresh input every call. Cached-read rates on the paid entries run from $0.006 per 1M (deepseek/deepseek-v4.1-flash) to $0.90 (openai/gpt-6-astra, first tier) — and on the models that carry a write rate, the write is the expensive side: $13.72 per 1M on anthropic/claude-fable-5.1 as the table displays it (billed at $13.725 before two-decimal rounding).
Cached-write pricing spans a wide range, which is why caching is not automatically a win. A prefix you write once and read many times amortises quickly; a prefix that changes on every request can cost more than it saves. The catalogue publishes the rates — what it does not publish is how a cache is declared, how long it lives, or the minimum size that qualifies. Treat the cached columns as the rate that applies when your provider reports a cache hit, and verify the behaviour on your own traffic before you build a cost model on top of it.
Then there is the label that trips up almost every reader of this table.
Two Gemini Flash rows show a 50% gap between the list price and the billed price — and that gap is Google’s, not ours. Both google/gemini-3.7-flash and google/gemini-3.8-flash list input at $1.50 and output at $7.50, with a billed rate of $0.75 and $3.75. Google’s own announcement for Gemini 3.8 Flash describes the model as available “at the same introductory price as 3.7 Flash at $0.75 per million input tokens and $3.75 per million output tokens” (Google, September 2026). Introductory means promotional: the $1.50/$7.50 column is the standard rate it reverts to. Modellix is a distribution layer, not the model maker — on those rows it adds nothing and discounts nothing. Google’s own footnote dates it: “Introductory price expires on December 31, 2026. Starting January 1, 2027, $1.50/1M input tokens and $7.50/1M output tokens will apply.” Budget against the standard rate, not the introductory one; for the full treatment of that row, see when the Gemini Flash promotional rate ends.
This is the general rule, and it is worth internalising before you read any gateway’s price table: a struck-through number can mean a gateway discount, a vendor promotional period, or a stale list price. Only the vendor’s own rate card settles which. Where the table’s two columns are identical, no discount of any kind is in play.
How the two-price cells and the discount column render on the live model table. Each cell shows the upstream list rate followed by the rate that is billed, and the discount column describes the gap between them — which, on the two Gemini Flash rows, is Google’s own schedule rather than a gateway discount. Table captured September 17, 2026.
Also on the “is it free” question, since it is the most common long-tail query in this keyword space: the catalogue currently has two entries priced at $0.00 across all four rate columns — modellix-ai/free-llm and zai/glm-4.7-flash. That is a pricing fact about two model entries. It is not a free tier: there is no automatic signup credit, and trial credit is a manual request. If a page tells you these models make API access free, it is describing something different from what the table says.
Choosing a coding LLM: the constraints, in the order that eliminates fastest
Coding is the highest-volume way people ask this question, and it is also where the ranking instinct does the most damage, because code workloads have constraints that dominate any quality ordering: repository-scale context, repeated context, and cost per iteration.
Apply the filters in this order.
1. Repo-scale context. An agent that reads files pulls a lot of tokens in. If your working set is large, the window decides before price does: the two 200,000-token entries — anthropic/claude-haiku-4.5 and modellix-ai/free-llm — accept roughly a fifth of what the 1.05M-window models do.
2. The tier boundary, because agent loops cross it. This is the coding-specific trap. An agent re-sends context on every turn; input accumulates; a single call crosses 272,000 input tokens; the rate doubles for that call. Coding sessions are exactly the workload where the subtle pricing mechanism from the previous section bites. Check your per-call input size, not your total session size — the tier follows the request.
3. Cached input, because coding context is repetitive. System prompts, tool definitions, and unchanged file contents repeat across turns. That is the ideal shape for cached reads, and the rate differences are large: $0.006 per 1M on deepseek/deepseek-v4.1-flash, $0.027 on zai/glm-5.3-flash, $0.36 on openai/gpt-5.6-sol, $0.45 on anthropic/claude-opus-5. On a long session this line item is not a rounding error.
4. Output price, because code generation is output-heavy. Code is long. A model at $0.47 output (qwen/qwen3.8-flash) and one at $45.00 output (anthropic/claude-fable-5.1) are two orders of magnitude apart on the same generated file.
5. Reasoning controls, because you may want to turn effort down. All three protocols expose a reasoning knob — reasoning_effort on Chat Completions, reasoning on Responses, thinking and output_config.effort on Messages — and support and allowed values depend on the model. Lower effort is the standard way to cut token overhead on routine edits, which is a cost lever the table does not show.
Notice what is absent from that list: a benchmark score. We are not ranking these models by coding ability here, and the reason is not modesty — it is that a benchmark number is only meaningful against the harness and version it was measured with, and we have not run that measurement ourselves. A ranked list you cannot reproduce is still a ranking; it is just not evidence. The constraints above you can reproduce, which is why they carry the decision on this page. (Once you have picked the model, the next question is where it runs — that is how each coding client gets its model source.)
Open weights, local runs, and one key for many models are three different decisions
This is where the most common category error in this keyword space happens, so it is worth separating carefully: “open source LLM”, “run it locally”, and “use one API key for many models” are not three answers to the same question. They are three different questions with different cost structures.
Open weights. Whether a model’s weights are downloadable is a decision made by the model maker, under a licence the model maker writes. The catalogue table lists model IDs, modalities, context, and prices — it does not publish a weights or licence field, so this table cannot tell you whether a given entry corresponds to downloadable weights. What you can see in the table is the boundary this page is about: calling a model through an API is a hosted-compute arrangement, while running weights yourself is a hardware and operations arrangement. Those are not substitutes for each other, and neither one is “the open-source option”. If licence terms decide your procurement, check them with the model’s publisher — not with a gateway’s price table.
Two families that dominate “open source LLM” discussions, Meta’s Llama and Mistral, are not in this catalogue at all. The catalogue offers no fine-tuning service either. It is a calling surface for 30 specific models, and knowing who is not on it is as useful as knowing who is.
Local runs. Quantised local inference — the GGUF, Ollama, and vLLM family of approaches — has one upfront cost (hardware you buy or rent) and near-zero marginal cost per token, with your own throughput ceiling. An API has zero upfront cost and a marginal cost on every token. The crossover depends on your volume and your utilisation; a cheap GPU is not cheaper than a $0.14 per 1M input rate if it sits idle, and a busy pipeline can make self-hosting look expensive against a $0.45 output rate. Nothing in the catalogue addresses this comparison, and a page that pretends to would be selling you a decision it cannot price.
One key for many models. Different question, and the one most often mislabelled. If your real problem is “I do not want five vendor accounts, five SDK versions, and five invoices for five models”, that is a routing question, and it is answered by a gateway. Concretely: one mdlx- key reaches all 30 text models plus the media catalogue, the model is selected by a provider/name string in the request body, and the same host serves all three protocol families — /v1 for OpenAI-compatible Chat Completions and Responses, the bare host for Anthropic-compatible Messages. Switching models is a field change, not a rewrite — which is the real answer to migration cost.
That last point is also where an honest comparison belongs, because the pricing structures genuinely differ. Another well-known gateway documents its approach as passing through the underlying provider’s pricing with no markup on inference, and charging a 5.5% fee on credit purchases, $0.80 minimum (OpenRouter’s FAQ). This catalogue’s structure is the other shape: several rows bill below the upstream list price, and top-ups carry a first-purchase discount rather than a purchase fee. Two different shapes, not two price points — one takes a cut when money goes in, the other adjusts on the token rate. Which is cheaper depends on your volume and whether a promotional rate on a specific model is still running, so run your own numbers; OpenRouter alternatives that bill per token covers that comparison. We are the second of the two shapes, so weigh that.
And here is where this page’s honesty has to bite: several things a procurement reviewer would ask are not published in the documentation we could read — there is no uptime or SLA commitment, no data-residency statement, and no statement about how long prompts are retained or whether they are used for training. The request-log interface does expose request and result payloads “when retained”, but the retention window itself is not specified. If any of those is a blocker for your organisation, ask support directly before you build. An unattended page should not quietly fill those gaps with confident prose.
What a real agent run actually bills
Per-million-token rates are hard to feel. Here is the same workload priced across the table, so the spread becomes concrete instead of theoretical.
The scenario: one coding-agent session, 2,000,000 input tokens across its turns and 40,000 output tokens, with no single request exceeding the tier boundary. Then the same session again, except 1,500,000 of those input tokens are a cached prefix. All rates are the billed (right-hand) values from the table above.
| Model ID | Input $/1M | Output $/1M | Session, no cache | Session with a 1.5M cached prefix |
|---|---|---|---|---|
anthropic/claude-fable-5.1 |
$9.00 | $45.00 | $19.80 | $6.64 |
anthropic/claude-sonnet-5 |
$1.80 | $9.00 | $3.96 | $1.53 |
openai/gpt-5.6-sol |
$3.60 | $18.00 | $7.92 | $3.06 |
openai/gpt-5.6-luna |
$0.18 | $1.08 | $0.40 | $0.16 |
xai/grok-4.6 |
$2.00 | $6.00 | $4.24 | $1.99 |
moonshot/kimi-k3 |
$3.00 | $15.00 | $6.60 | $2.55 |
google/gemini-3.8-flash |
$0.75 | $3.75 | $1.65 | $0.64 |
deepseek/deepseek-v4.1-flash |
$0.30 | $1.20 | $0.65 | $0.21 |
qwen/qwen3.8-flash |
$0.16 | $0.47 | $0.34 | $0.12 |
zai/glm-5.3-flash |
$0.14 | $0.45 | $0.29 | $0.13 |
Three readings from that table are worth more than the numbers themselves.
Caching changes the ranking more than it changes the total. On the same session, anthropic/claude-fable-5.1 bills $19.80 uncached and $6.64 with a cached prefix — a 66% reduction. openai/gpt-5.6-sol falls from $7.92 to $3.06, also about 61%. A model that looks unaffordable on the headline rate can be competitive once most of its input is a repeat, which is precisely the profile of an agent that re-reads a codebase.
Tier crossings are invisible in session totals until they are not. The numbers above assume no single call crosses a boundary. Allow one 300,000-token call into openai/gpt-5.6-sol and that call bills at $7.20 input and $27.00 output instead of $3.60 and $18.00. Nothing in the session total warns you; the request just costs double.
Small workloads are dominated by output price. A chat-shaped app doing 5,000 calls at 1,500 input tokens and 400 output tokens each bills $1.95 on zai/glm-5.3-flash, $3.51 on openai/gpt-5.6-luna, $13.12 on google/gemini-3.7-flash, $31.50 on anthropic/claude-sonnet-5, and $157.50 on anthropic/claude-fable-5.1. If your traffic is request-heavy and answer-light, optimise output rate first; if it is document-heavy, optimise input and cache.
Treat all five of those as arithmetic on published rates, not as predictions of your invoice — and reproduce them with the per-request cost field described below, or with the worked example in what one agent run actually bills, from the logs. These figures use the two-decimal rates as the table displays them; recomputing from unrounded values can move a total by a cent.
Verify every number on this page yourself, in about four commands
A price table is only as good as your ability to re-check it. Everything on this page came from a public endpoint or a public page, and none of it needs our cooperation to confirm.
One: pull the current model list for the text gateway. GET /v1/models on the LLM host returns an OpenAI-compatible list whose data[].id values are the provider/name strings used above, plus a display name, provider name, and series name. It runs on the query rate limit, which the documentation states is tracked separately from the inference quota — so a listing call does not eat your request budget. The same page also publishes stable aliases for each model family — ~openai/gpt-latest currently routes to openai/gpt-6-astra, ~google/gemini-flash-latest to google/gemini-3.8-flash — so a version bump on the vendor’s side is a routing change on the gateway’s side rather than an edit on yours. Aliases bill at whatever their current target costs, and the documentation warns against using them for evaluations or regression tests; pin a concrete model ID for those. If the phrase “OpenAI-compatible” is doing more work in that sentence than you would like, an OpenAI-compatible API is a protocol contract rather than a compatibility badge.
Two: inspect a model’s request contract without spending anything. The modellix-cli package ships a schema command that is public and needs no API key (source on GitHub, published on npm as modellix-cli):
1 | npm install --global modellix-cli@latest |
That prints the full public schema for the slug — servers, request body, examples, and response shapes. --output human summarises the contract, and --quiet prints only the inference URL for scripting. The reason this command exists is worth repeating, because it is the difference between two kinds of page: it was added so that agents and scripts read the schema directly, “instead of guessing fields from documentation pages.” Use it the same way. Listing slugs requires an API key (model list --output slugs), but once you have a slug, no credentials are needed to inspect its contract.
The whole point of the command: request body, field constraints, and examples pulled from the model’s live schema, with no credentials in the environment. Transcript captured September 17, 2026 from modellix-cli 0.0.10.
See the Modellix API reference
Check the request fields, error codes and log endpoints behind the numbers on this page.
View DocsThree: batch a comparison instead of hand-running it. modellix-cli model batch takes a JSONL file — one JSON object per line — so the same prompt can be submitted against several models in one pass:
1 | modellix-cli model batch runs.jsonl --max-tasks 6 --concurrency 3 --wait |
Batch submission requires either --max-tasks or an explicit -y, because every line can create a paid task; concurrency is limited to 1–10 and the local ceiling is 1,000 tasks. The batch file is parsed and validated before the first POST, so a malformed line aborts the run rather than halfway through a paid batch, so check the IDs first. That is the cheapest way to generate your own comparison data rather than borrowing someone else’s — including ours.
Four: reconcile the invoice against per-request logs. GET /v1/logs returns each request with prompt_tokens, completion_tokens, cached_tokens, and a cost field in the console’s billing unit, scoped to your team, with page_size up to 100 and a maximum window of 30 days. If a model’s billed cost surprises you, this is where you find out whether the cause was a tier crossing or a cache miss; agent observability for gateway request logs covers what this layer can and cannot show you. Two rate-limit classes are worth knowing here as well: inference requests and query requests (model listing, log listing) are separately metered, and 429 responses distinguish rate limits from temporarily unavailable models, the latter sometimes carrying Retry-After.
Two honest gaps while you are in there. The gateway documentation does not describe structured output, JSON mode, or a tool-calling field — we could not confirm either support or absence, so do not plan around it without testing. And the catalogue’s two $0.00 entries have no documentation page describing limits or availability, so treat them as priced entries you must probe, not as a supported free plan.
Three things this page does not measure and the catalogue does not publish: accuracy or hallucination rates, latency or time to first token, and a moderation or safety policy. If any of those decides your choice, none of the numbers here will help you — and a page that supplied them without a reproducible method would be guessing on your behalf.
A note on the media side, since one key covers both. If your pipeline also generates images or video, media model invocation paths no longer carry the /async suffix — POST /api/v1/alibaba/qwen-image-3.0-pro instead of the older .../async form — and the older endpoints continue to work. The change was announced as backwards compatible, so existing integrations are unaffected and there is no migration deadline to plan for. Media discovery uses a different endpoint from the text list: GET /api/v1/models returns each model’s slug, type, name, series_name, docs_url, a dynamic description, and a price object with a unit such as $/img, which is the fastest way to find a slug for model get-schema.
How to choose: match the constraint you actually have
There is no first place on this page, so here is the routing instead.
| If your binding constraint is… | Start with | Because |
|---|---|---|
| An image, video, or audio input | The 24 / 8 / 4 modality groups in filter one | Modality is a hard filter; nothing else matters until it passes |
| A very large single request | zai/glm-5.3-flash (1,310,720) or the 1.05M group |
Window caps the request before price does |
| Predictable cost on unpredictable input | Any flat-rate entry (22 of 30) | No tier boundary means no doubling surprise |
| A long session that re-sends context | Models with the lowest cached-read rate | Caching moved a $19.80 session to $6.64 in the example above |
| Generation-heavy output | The low output-rate entries ($0.45–$1.20) | Output rates span about 100×, first-tier input about 64× |
| Cheapest possible iteration on routine edits | zai/glm-5.3-flash, qwen/qwen3.8-flash, openai/gpt-5.6-luna |
Lowest input, output and tier-1 rates in the table |
| Not wanting five vendor accounts | A gateway with a provider/name model field |
Model switching becomes a body field, not a rewrite |
| Reproducible comparison data you own | modellix-cli model batch over your own prompt set |
Your traffic, your prompts, your numbers |
A practical sequence that costs little: pick the three candidates that survive the filters, run your own ten-prompt set against each, and compare the per-request cost figures from the logs rather than the headline rates. Gate the decision on your measured numbers — including a deliberately long request, so you see whether a tier boundary is waiting for you. And if the fork you are actually weighing is a self-hosted proxy against a hosted gateway, rather than one gateway against another, self-hosting an LLM proxy versus a hosted gateway takes that one head-on.
And keep the table dated in your own notes. Rates in this catalogue moved within the month we looked at it: the entry count went from the 28 rows captured on September 3 to the 30 rows above, and promotional schedules on individual models change on the vendor’s calendar, not the gateway’s. A price comparison without a capture date is a rumour.
Nobody can tell you which LLM is best. The catalogue will tell you which one fits, and the commands above let you check that it is still true next month.
Frequently Asked Questions About the Best LLMs
Is there a best LLM overall in 2026?
No, and the question cannot be answered as phrased. “Best” requires a workload: the same model can be the right choice for a large-context retrieval pipeline and the wrong one for high-volume short chat. What is answerable — and what the table above answers — is which models pass your constraints on input modality, context size, and price mechanics.
What are the top-ranked LLM models right now?
That is what live leaderboards measure, and they do it better than a static page can. Two caveats worth carrying: leaderboards rank on their own metrics (price, throughput, preference votes, or routed usage), which are not your metrics; and their model inventories move weekly. Use them for the frontier question, and use a price-and-constraint table like this one for the decision that follows.
Is there a better LLM than ChatGPT?
This question blends a product and a model, which is why the answers never agree. ChatGPT is an application built on top of models; the models are also callable individually. Whether something is “better than ChatGPT” depends on whether you are comparing chat experiences or per-token economics — for the second, the relevant comparison is between the model IDs in the table, not between consumer apps.
Which LLM is best for coding, and what makes a best coding LLM?
The one that survives your constraints, tested in this order: context window large enough for your working set, a per-call input size that stays inside the first pricing tier, a low cached-read rate for repeated context, a low output rate for generated code, and reasoning controls you can dial down for routine edits. Ranking by benchmark score instead usually optimises for a harness you are not running.
Are any of these LLMs free to use through an API?
There is no automatic signup credit (discontinued in August 2026); trial credit is a manual request. Two catalogue entries are priced at $0.00 per million tokens across all four rate columns, but neither has a documentation page describing limits or availability, so test capacity before you design around them. Pay-as-you-go with no monthly fee is the actual pricing model — not a free tier.
Which AI is best for LLM?
Read as written, this question asks which assistant is best for working with language models; read as intended, it usually means “which model should I pick”. For the second reading, the answer is the filter sequence on this page — modality, then context, then pricing mode — because those three attributes are published for every model and constrain the choice harder than any quality ordering.
Can I switch between these models without rewriting my integration?
Yes — the model is a field in the request body, not a rewrite. The mechanics, and the protocol surface they sit on, are in the “one key for many models” section above.
Which NSFW LLM is the best?
This question appears in the search results for this term, and it is not a dimension the catalogue or this page evaluates. The 30 models listed here are general-purpose text models with published input modalities and token pricing, and no uncensored or unrestricted variant is among them. If that is your requirement, this page will not help you; if your requirement is cost, context, or modality fit, everything above applies.
Model IDs, context windows, modalities and rates were captured from the live Modellix model table on September 17, 2026; promotional rates set by model vendors change on their own schedule, and gateway catalogues change weekly. Re-check every figure against the current vendor table before you commit spend. Access these 30 language models — plus the image and video catalogue — through a single API key at modellix.ai.