Direct
Direct upstream connection — best when you need native behavior and the full context window.
| Input context | Input | Output | Cache read | Cache write |
|---|---|---|---|---|
| ≤ 272K | 0.10/M | 0.50/M | 0.01/M | 0.13/M |
| > 272K | 0.20/M | 0.75/M | 0.02/M | 0.25/M |

gpt-6-lunaGPT-6 Luna is an economy-tier model released by OpenAI on September 22, 2026, on the same day as GPT-6 Sol. It is positioned for clearly defined, high-volume tasks—document summarization, information extraction, and rapid Q&A.
The same model is available through multiple service channels — choose based on latency, reliability and cost.
Prices in $ / 1M tokensprovider field to the request body, for example "provider": { "channel": "direct" }. Valid values are direct / stable / economical; omit it to use the default channel.Direct upstream connection — best when you need native behavior and the full context window.
| Input context | Input | Output | Cache read | Cache write |
|---|---|---|---|---|
| ≤ 272K | 0.10/M | 0.50/M | 0.01/M | 0.13/M |
| > 272K | 0.20/M | 0.75/M | 0.02/M | 0.25/M |
GPT-6 Luna is an economy-tier model released by OpenAI on September 22, 2026, the same day as GPT-6 Sol. It is positioned for well-defined, high-volume tasks—document summarization, information extraction, and fast Q&A. It is also the GPT-6 model available to ChatGPT Free and Go users.
The core change in this generation is cost: the official input price is half that of GPT-5.6 Luna, and the output price is reduced by 58.3%. Artificial Analysis measured the per-task cost of running the full Intelligence Index dropping from $0.18 to $0.07, a decrease of about 60%. The Intelligence Index score is 37, on par with GPT-5.6 Luna.
Most notable is factuality: OpenAI says that at higher reasoning effort, Luna's factuality can match GPT-5.6 Sol at about one-hundredth of its cost. Third-party data points in the same direction—the AA-Omniscience hallucination rate dropped from 93% to 77%.
SeaWhale AI provides GPT-6 Luna through an OpenAI-compatible API, supporting tool calling, reasoning effort control, streaming output, structured outputs, and image-and-text input.
Get an API key · Model ID:
gpt-6-luna
This is OpenAI's clear positioning for Luna: well-defined, repeatable, high-frequency tasks. Combined with structured outputs and the Batch API, work such as document summarization, field extraction, and classification and labeling can be scaled out at a very low unit cost. AA measured an output speed of about 157 tokens/second, which is relatively fast among reasoning models at the same price point.
Luna shows clear factuality improvements in this generation: the AA-Omniscience score improved from -10 to 1, and the hallucination rate dropped from 93% to 77%. OpenAI's internal factuality evaluation shows that at higher reasoning effort, Luna can match GPT-5.6 Sol's level. For large volumes of simple Q&A, this means there is no need to upgrade to the flagship tier just for reliability.
AutomationBench-AA improved from 50.2% to 53.2%, and Terminal-Bench 4.0 improved from 11.6% to 12.6%. In multi-agent architectures, Luna is suited to subtasks such as retrieval, filtering, and simple tool calling, while leaving parts that require deep reasoning to Sol.
The official DeepSWE v1.1 score at the max setting is 66.6%, only 2.2 percentage points behind GPT-6 Sol's 68.8%. However, the AA Coding Agent Index is 41, 2 points lower than GPT-5.6 Luna. It is suitable for code completion, simple fixes, and batch code processing; for complex engineering tasks, Sol is still recommended.
| Use Case | Description |
|---|---|
| Document summarization | High throughput, low unit cost; the official top recommended use case |
| Information extraction | Generate structured JSON from unstructured text |
| Fast Q&A | Hallucination rate down to 77%, suitable for high-frequency simple Q&A |
| Content classification and moderation | Intent recognition, labeling, compliance checks |
| Sub-agents | Low-cost execution units in multi-agent architectures |
| Offline batch processing | Process large-scale corpora with the Batch API |
| Capability | GPT-6 Luna | GPT-5.6 Luna | GPT-6 Sol |
|---|---|---|---|
| Model ID | gpt-6-luna |
gpt-5.6-luna |
gpt-6-sol |
| AA Intelligence Index (v4.3) | 37 | 37 | 48 |
| AA Coding Agent Index | 41 | 43 | 57 |
DeepSWE v1.1 (official, max) |
66.6% | — | 68.8% |
| AA-Omniscience hallucination rate | 77% | 93% | 60% |
| AA per-task cost | $0.07 | $0.18 | $1.06 |
| Context window | 1.05M tokens | — | 1.05M tokens |
| Max output | 128K tokens | — | 128K tokens |
| Knowledge cutoff | May 18, 2026 | — | April 20, 2026 |
| Positioning | High-throughput economy tier | Previous-generation economy tier | Cost-effective primary tier |
Actual billing is subject to the real-time pricing card at the top of the page.
When was GPT-6 Luna released? September 22, 2026, released the same day as GPT-6 Sol. It launched in the API, Codex, and ChatGPT that day, and was also made available to ChatGPT Free and Go users.
What has been upgraded compared with GPT-5.6 Luna? Mainly cost and factuality: per-task cost dropped by about 60%, the hallucination rate fell from 93% to 77%, and business automation (AutomationBench-AA) improved slightly by 3 percentage points. The overall Intelligence Index is unchanged; both are 37.
Are there any areas where it has regressed? Yes. The AA Coding Agent Index fell from 43 to 41, and SWE-Atlas-QnA fell from 49% to 44%; knowledge-work evaluations also regressed—GDPval-AA v2.1 fell from about 1443 Elo to 1367, and AA-Briefcase fell from 1345 to 1299. It is also more verbose: completing the same batch of tasks uses about 23% more output tokens, which partially eats into the price reduction. When a single prompt exceeds 272K tokens, input is billed at 2x and output at 1.5x.
What are the context and output limits? Context window is 1.05M tokens, with a maximum input of 922K and a maximum output of 128K tokens. Input supports text and images; output is plain text. The knowledge cutoff is May 18, 2026, about one month later than Sol.
How should reasoning effort be set?
reasoning_effort supports six levels: none, low, medium (default), high, xhigh, and max. For simple tasks such as classification and extraction, none or low is recommended for the lowest latency and cost; raise it when factuality close to flagship level is needed. Note: when using function calling through the Chat Completions API, the official requirement is to set reasoning_effort to none; if you need reasoning and tool calling enabled at the same time, use the Responses API.
How to choose between Luna and Sol? Use Luna for well-defined, repeatable, high-frequency tasks; use Sol for multi-step reasoning, complex coding, or business workflows that span multiple tools. The two have identical APIs and parameters, so you can switch at runtime based on request complexity.
gpt-6-lunahttps://api.haijingai.com/v2/"provider": { "channel": "direct" }SeaWhale AI is compatible with the OpenAI API protocol, so you can call it with the OpenAI SDK or plain HTTP requests. Streaming is enabled by default.
curl https://api.haijingai.com/v2/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <API_KEY>" \
-d '{
"model": "gpt-6-luna",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"}
],
"provider": { "channel": "direct" },
"stream": true
}'
# provider is optional — remove this line to use the default channelfrom openai import OpenAI
client = OpenAI(
base_url="https://api.haijingai.com/v2",
api_key="<API_KEY>",
)
stream = client.chat.completions.create(
model="gpt-6-luna",
messages=[
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "Hello!"},
],
stream=True,
# Optional: pick a service channel; omit to use the default
extra_body={"provider": {"channel": "direct"}},
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)import OpenAI from 'openai'
const client = new OpenAI({
baseURL: 'https://api.haijingai.com/v2',
apiKey: '<API_KEY>',
})
const stream = await client.chat.completions.create({
model: 'gpt-6-luna',
messages: [
{ role: 'system', content: 'You are a helpful assistant.' },
{ role: 'user', content: 'Hello!' },
],
stream: true,
// Optional: pick a service channel; omit to use the default
// @ts-expect-error provider is a SeaWhale AI extension, not in the OpenAI SDK types
provider: { channel: 'direct' },
})
for await (const chunk of stream) {
process.stdout.write(chunk.choices[0]?.delta?.content ?? '')
}