The frontier race gets the headlines. The price war at the bottom of the menu is what changes your bill. On September 22, OpenAI released GPT-6 Luna at $0.10 per million input tokens and $0.50 per million output tokens, a tenth of what Claude Haiku 4.5 costs. Most real-world AI work is classifying, extracting, routing, summarizing and answering simple questions, and the cost of doing it just dropped by an order of magnitude.
Our verdict: GPT-6 Luna is the cheapest AI model worth using in 2026 and the new default for high-volume work. Gemini 3.8 Flash is the step-up pick when you need real reasoning and agent behavior without frontier prices. Claude Haiku 4.5 is now the most expensive budget model and makes sense mainly for teams already built on Anthropic.
The Short Version
Cheapest model worth using: GPT-6 Luna, $0.10 / $0.50.
Best budget model for agents and coding: Gemini 3.8 Flash, $0.75 / $3.75 on promotional pricing through December 31.
Budget pick for Anthropic users: Claude Haiku 4.5, $1 / $5.
What Each One Costs at Real Volume
| Model | Released | Input / 1M | Output / 1M | Monthly bill* | Best at |
|---|---|---|---|---|---|
| GPT-6 Luna | Sep 22, 2026 | $0.10 | $0.50 | $700 | Volume, extraction, routing |
| Gemini 3.8 Flash | Sep 2, 2026 | $0.75 (promo) | $3.75 (promo) | $5,250 | Reasoning, agents, coding |
| Claude Haiku 4.5 | 2025 | $1.00 | $5.00 | $7,000 | Anthropic stacks, writing quality |
*Example: a support or document pipeline processing 2 billion input tokens and 1 billion output tokens a month. At that scale, choosing Luna over Haiku saves $6,300 a month, or about $75,000 a year. That's why this decision deserves ten minutes of testing.
GPT-6 Luna: The New Default for Volume
Luna is OpenAI's smallest GPT-6 model, priced to win the high-volume market outright. It's half the cost of the previous small OpenAI model and comes with OpenAI's mature tooling: function calling, structured outputs and the largest ecosystem of libraries and integrations. Use it for tagging and classification, pulling fields out of documents, routing requests to bigger models, first-pass summaries and simple customer chat.
The catch: small models still stumble on long chains of reasoning and unusual edge cases. Add a quality check to anything customer-facing and send the hard cases to a bigger model.
Gemini 3.8 Flash: The Smart Middle
Google released 3.8 Flash on September 2, just three weeks after 3.7 Flash, and it's a clear step up in software engineering, agent workflows and multi-step reasoning. At $0.75/$3.75 on promotional pricing through December 31, it costs more than Luna but far less than any frontier model. It's the right pick for lightweight agents, coding helpers and anything that has to think before it answers. Google also released a Cyber variant that posted frontier-level results on finding security vulnerabilities autonomously.
The catch: the promotional price ends December 31. Budget for a higher 2027 rate.
Claude Haiku 4.5: Good, but Now Expensive
Haiku 4.5 is still fast, reliable and a noticeably better writer than most small models. But at $1/$5 it costs ten times Luna and more than Gemini 3.8 Flash's promo price. If your product already runs on Claude Code, Anthropic's SDK or Claude Skills, consistency may be worth paying for. If you're starting fresh, start somewhere else.
When to Stop Being Cheap
Budget models should handle most of your requests, not all of them. When a wrong answer is expensive (hard coding, long agent runs, legal or financial analysis, anything where quality is the product), send it to a frontier model. Right now that's Claude Opus 5.5 or GPT-6 Sol, compared in our September frontier ranking. The best setups use a cheap model to sort requests and a strong model for the hard 10%.
The Verdict
GPT-6 Luna is the cheapest AI model worth using in 2026 and the new default for high-volume work at $0.10/$0.50. Gemini 3.8 Flash is the better budget model for reasoning and agents. Claude Haiku 4.5 is the right choice mainly if you're already committed to Anthropic.
FAQ
What is the cheapest good AI model in 2026?
GPT-6 Luna, released September 22, 2026, at $0.10 per million input tokens and $0.50 per million output tokens. That's one tenth of Claude Haiku 4.5 and far below Gemini 3.8 Flash's promotional $0.75/$3.75.
Is Gemini 3.8 Flash better than GPT-6 Luna?
For reasoning, yes. Gemini 3.8 Flash is stronger at software engineering, agent workflows and multi-step reasoning, and it's priced as a step-up tier at $0.75/$3.75 through December 31, 2026. GPT-6 Luna is the better pick when volume and cost matter most.
Is Claude Haiku 4.5 still worth it?
Mostly for teams already built on Anthropic's API and tools. At $1/$5 per million tokens it now costs ten times GPT-6 Luna. If you're choosing fresh, start with Luna or Gemini 3.8 Flash.
How much can I save by switching to GPT-6 Luna?
At 2 billion input and 1 billion output tokens a month, GPT-6 Luna costs about $700, compared with about $7,000 for Claude Haiku 4.5 and $5,250 for Gemini 3.8 Flash at promotional prices.
When should I pay for a frontier model instead?
When a wrong answer is expensive: hard coding, long multi-step agent tasks, legal or financial analysis and anything customer-facing where quality is the product. Route those to Claude Opus 5.5 or GPT-6 Sol and keep budget models for sorting, extraction and simple chat.