Cost Savings On GPT‑6 Sol And Luna: OpenAI Keeps Benchmark Performance Steady
AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Cost Savings On GPT‑6 Sol And Luna: OpenAI Keeps Benchmark Performance Steady on ThorstenMeyerAI.com

Buying for a business?Offer from Amazon

Get business pricing on monitors, keyboards and dev gear

  • Business-only prices and quantity discounts
  • Tax-exempt purchasing
  • Multiple users, one account, clear invoices
As an affiliate, we earn on qualifying purchases.

TL;DR

OpenAI has launched GPT‑6 Sol and Luna models at 50% lower prices than their GPT‑5.6 predecessors, achieving significant cost savings while maintaining comparable performance levels. This move aims to expand AI accessibility and application viability.

OpenAI has introduced GPT‑6 Sol and GPT‑6 Luna models at approximately half the cost of their GPT‑5.6 predecessors, without sacrificing benchmark performance. The models, released on September 22, 2026, focus on delivering cost-efficient AI solutions for a broad range of applications, emphasizing affordability for businesses and developers aiming to embed advanced AI capabilities into their workflows.

Both GPT‑6 Sol and Luna benefit from enhanced caching and inference techniques, enabling lower operational costs. GPT‑6 Sol is priced at $2.00 per 1 million input tokens and $10.00 per 1 million output tokens, down from $4 and $20 respectively for GPT‑5.6. Luna is even more affordable, at $0.10 per 1 million input tokens and $0.50 per 1 million output tokens, compared to $0.20 and $1.20 previously.

Independent analysis by Artificial Analysis indicates that while costs have halved, the models’ composite scores on intelligence benchmarks remain roughly stable, with Sol scoring 48 and Luna 37 on their respective indexes. Notably, the models show improvements in hallucination reduction, with Sol decreasing its hallucination rate from 92% to 60%, and Luna from 93% to 77%, partly by increasing refusal rates for uncertain responses.

However, some regressions were observed in knowledge-work benchmarks, with both models scoring lower on economic and multi-week task evaluations, attributed to a shift toward less detailed output. These trade-offs highlight the balance between cost, accuracy, and output quality, which users need to consider based on their specific needs.

At a glance
updateWhen: announced September 22, 2026
The developmentOpenAI announced the release of GPT‑6 Sol and Luna on September 22, 2026, offering the same or better performance at half the previous costs, driven by improved caching and inference efficiencies.

GPT‑6 Sol and Luna: half the price, about the same intelligence

OpenAI’s September 22, 2026 release doesn’t raise the ceiling. It lowers the cost of everything below it, which changes what’s worth automating.

GPT‑6 Sol
$4 / $20 → $2 / $10
GPT‑6 Luna
$0.20 / $1.20 → $0.10 / $0.50

Per 1M input / output tokens. Cached input reads keep the 90% discount.

Cost per task, halved

Measured by Artificial Analysis as the weighted cost of one Intelligence Index task, at max effort.

GPT‑5.6 Sol
$1.99
GPT‑6 Sol
$1.06
GPT‑5.6 Luna
$0.18
GPT‑6 Luna
$0.07

The effort dial moves cost more than the model choice

Model and effortIntelligence IndexCost per task
GPT‑6 Sol (max)48$1.06
GPT‑6 Sol (low)34$0.13
GPT‑6 Luna (max)37$0.07
GPT‑6 Luna (low)21$0.0045
GPT‑6 Luna (non‑reasoning)18$0.01

Sol at low effort keeps about 70% of its max score for roughly an eighth of the cost, because it writes far fewer reasoning tokens. For reference, Claude Opus 5.5 leads the same index at 58.

What got better, and what got worse

Better

  • Hallucination rate on AA‑Omniscience: Sol 92% → 60%, Luna 93% → 77%
  • Coding Agent Index: Sol 57, up 2 points, at ~50% lower cost per task
  • OpenAI reports about half as many factual mistakes for Sol as its predecessor
  • Higher cache hit rates; GitHub reports over 50% fewer prompt tokens needing fresh processing

Sol gets there partly by declining more: it attempts 83% of questions vs 99%, and accuracy falls 59% → 54%.

Worse

  • GDPval‑AA v2.1: Sol down ~100 Elo, Luna down ~75
  • AA‑Briefcase v1.1: Luna down ~45 Elo
  • Coding Agent Index: Luna 41, down 2 points
  • Both models write more output tokens per task than their predecessors

Reviewers attribute the drops to weaker presentation and deliverables that omit required elements.

What to do about it

Already on GPT‑5.6 Sol or Luna? The move is mostly a price cut. Re‑test first if your output is a document someone reads, not data a system consumes.
Shelved an automation on cost? Token prices halved and the effort dial adds another order of magnitude. Re‑run the business case.
Choosing between labs? The question is no longer which model is smartest, but which clears your quality bar at the lowest cost per task.
ThorstenMeyerAI.comSources: OpenAI (pricing, vendor benchmarks) and Artificial Analysis (independent evaluation and model pages). Figures as of 23 September 2026.

Impact of Cost-Effective AI on Business and Development

The release of GPT‑6 Sol and Luna at significantly reduced costs is poised to broaden AI adoption across industries. By maintaining performance benchmarks while lowering expenses, OpenAI enables companies to automate more tasks, deploy AI in cost-sensitive applications, and scale AI solutions more efficiently. This shift could accelerate AI integration in customer service, research, and content generation, making advanced language models accessible to a wider range of users and use cases.

Amazon

AI language model API subscription

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on OpenAI’s Model Pricing and Performance Trends

Prior to this release, OpenAI’s GPT‑5.6 models set a high-performance standard but at a relatively high cost, limiting their deployment in cost-sensitive scenarios. The company has emphasized ongoing efficiency improvements, including caching and inference optimizations, as key drivers behind the cost reductions. The September 2026 launch of GPT‑6 Sol and Luna marks a strategic shift toward prioritizing affordability without sacrificing the core capabilities that make these models valuable for enterprise and developer use.

Earlier, OpenAI’s models faced criticism for high operational costs, which constrained their widespread adoption. The new models aim to address this by passing savings onto users, thereby expanding the potential user base and use case diversity.

Amazon

cost-effective AI chatbot development tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unresolved Questions About Model Trade-offs and Long-term Performance

While initial benchmarks are promising, it remains unclear how these models will perform across diverse real-world tasks over extended periods. The observed regressions in knowledge-work benchmarks suggest potential limitations in detailed output quality, especially for complex or nuanced applications. Additionally, the long-term impact of increased refusal rates on user workflows and overall accuracy is still being evaluated. Further testing is needed to confirm whether these models can sustain their performance in varied operational environments.

Amazon

AI inference optimization hardware

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Adoption and Performance Validation

OpenAI is expected to continue monitoring the models’ performance across different industries and use cases. Users should conduct their own testing, especially for tasks requiring high detail or accuracy, before full deployment. Future updates may include further tuning to address current regressions and to optimize the balance between cost and output quality. Meanwhile, OpenAI will likely expand its caching tools and diagnostics to help developers better manage model efficiency and costs.

Amazon

AI model caching solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

How do GPT‑6 Sol and Luna compare to previous models in terms of cost?

They are priced at approximately half the cost of GPT‑5.6 models, with Sol at $2.00 per million input tokens and Luna at $0.10, significantly reducing operational expenses.

Do the new models maintain the same performance as GPT‑5.6?

Benchmark scores suggest performance remains roughly comparable, although some knowledge-work evaluations showed regressions, likely due to changes in output presentation and detail.

What are the main improvements that enable cost savings?

Enhanced caching, inference optimizations, and better reuse of context data contribute to lower costs without sacrificing core capabilities.

Are there any trade-offs with the new models?

Yes, some regressions in detailed output quality and knowledge-work performance have been observed, alongside increased refusal rates for uncertain responses.

What should users do before deploying these models in critical applications?

Users should conduct thorough testing to ensure the models meet their quality and accuracy needs, especially for complex or detailed tasks.

Source: ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Unveiling A Buried File Through AI Agent Testing

AI testing reveals how deep document reading impacts commercial success, exposing a buried fact that determined a €55,000 deal.

Kill-Switch-Proof: How to Build So Washington Can’t Take Your AI Stack Down

Discover the strategies to make AI infrastructure resilient against government shutdowns, including dependency mapping and open-weight models.

28 Startups Using AI To Transform The Energy Sector

Twenty-eight startups are applying AI to innovate energy production, management, and distribution, signaling a major shift in the industry.

SpaceXAI’s Grok 4.6 Sets New Benchmark In AI With GPT-5.6 And Claude Fable 5-Level Intelligence

SpaceXAI announces Grok 4.6, claiming it matches GPT-5.6 and Claude Fable 5 in intelligence, but lacks independent verification or benchmark data.