The rates that drive most bills
OpenAI prices every text model per million tokens, split into input, cached input, cache writes, and output. The three current flagship models span two orders of magnitude. At standard short-context rates on 2026-10-07, gpt-6-luna is $0.10 input, $0.01 cached input, $0.125 cache writes, and $0.50 output. gpt-6.1-sol is $2.00, $0.10, $2.50, and $10.00. gpt-6-astra is $10.00, $1.00, $12.50, and $50.00.
Each model has a second, higher column for long context. For gpt-6.1-sol, long-context input is $4.00 and output is $15.00, so a long prompt doubles the input rate and raises the output rate by half. The pricing table does not print the cutoff for the gpt-6 models; the older gpt-5.5 and gpt-5.4 rows label their short-context tier as under 272K tokens. Check the context length of your largest requests before you estimate from the short-context column.
- gpt-6-luna: $0.10 input, $0.50 output per million tokens.
- gpt-6.1-sol: $2.00 input, $10.00 output per million tokens.
- gpt-6-astra: $10.00 input, $50.00 output per million tokens.
- Older models stay on the price list, for example gpt-5-mini at $0.25 input and $2.00 output.
Four workloads, worked out
A support chat turn with a 1,500-token cached system prompt, 500 new input tokens, and 300 output tokens costs about $0.22 per 1,000 requests on gpt-6-luna, $4.15 on gpt-6.1-sol, and $21.50 on gpt-6-astra. Output is the largest share on every model: on gpt-6.1-sol, 300 output tokens cost $0.003 per request while 2,000 input tokens cost $0.00115. The first request that writes the cache pays the cache-write rate instead of the cached rate.
An overnight document job of 10,000 documents, each 4,000 input tokens and 500 output tokens on gpt-6.1-sol, is 40 million input and 5 million output tokens. At standard rates that is $130. Through the Batch API it is $65, because batch rates for these models are exactly half of standard.
A research agent on gpt-6-luna that makes 5 web search calls, reads 30,000 input tokens, and writes 2,000 output tokens spends about $0.004 on tokens and $0.05 on search call fees, since web search is $10.00 per 1,000 calls. On a cheap model, tool fees can be more than ten times the token cost, so the model choice barely moves the total. Search content tokens are billed at model rates on top of the call fee for most models.
A latency-sensitive endpoint on Fast mode pays twice the standard rate: gpt-6.1-sol is $4.00 input and $20.00 output. Ultrafast, offered for gpt-6-astra, is $60.00 input and $300.00 output, six times standard.
Multipliers that push a bill above the token math
A per-request estimate built only from input and output rates usually comes in low. The pricing page lists several charges that sit outside the main columns, and each one is easy to miss in a spreadsheet.
- Long context: higher rates on every token in requests above the short-context tier.
- Cache writes: priced above uncached input, $2.50 versus $2.00 per million on gpt-6.1-sol.
- Regional processing (data residency) and FedRAMP endpoints: a 10% uplift for models released on or after March 5, 2026.
- Hosted Shell and Code Interpreter containers: $0.03 to $1.92 per 20-minute session depending on memory size.
- File search storage: $0.10 per GB per day after the first free GB, plus $2.50 per 1,000 tool calls.
- Fast and Ultrafast modes: 2x and 6x standard rates on supported models.
Cutting the cost without changing the product
The largest lever is model choice per request, because the gap between gpt-6-luna and gpt-6-astra is 100x on the chat example above. Routing simple requests to a small model and reserving the large one for hard cases cuts more than any discount tier. The second lever is the service tier. OpenAI's Batch guide offers a 50% discount with a 24-hour completion window, and its Flex processing guide prices synchronous requests at Batch rates in exchange for slower responses. Flex can return 429 Resource Unavailable when capacity is short, and OpenAI says those requests are not charged.
The third lever is caching. Cached input on gpt-6.1-sol is $0.10 per million, one twentieth of the uncached rate, so a long, stable system prompt placed at the start of every request pays for itself quickly. To see which of these levers matters for your traffic, group your bill by model and project first. FrugalAI does this across providers and enforces budgets per key, but the same analysis works from OpenAI's own usage and cost exports.
Frequently asked questions
How much does the OpenAI API cost per request?
It depends on the model and token counts. As of 2026-10-07, a chat request with 2,000 input tokens (1,500 of them cached) and 300 output tokens costs about $0.00415 on gpt-6.1-sol, $0.000215 on gpt-6-luna, and $0.0215 on gpt-6-astra at standard rates.
Is the OpenAI Batch API really half price?
Yes for the listed models. OpenAI's Batch guide states a 50% discount compared to synchronous APIs, and the batch column on the pricing page is half the standard column, for example $1.00 input and $5.00 output per million tokens on gpt-6.1-sol. Batches complete within 24 hours.
Do the Responses API and Chat Completions API cost different amounts?
No. OpenAI's pricing page says the Responses, Chat Completions, Realtime, Batch, and Assistants APIs are not priced separately. Tokens are billed at the chosen model's rates, and built-in tools such as web search and file search add their own fees.
Sources and further reading
- OpenAI API pricing (verified 2026-10-07)
- OpenAI Batch API guide (verified 2026-10-07)
- OpenAI Flex processing guide (verified 2026-10-07)
FrugalAI uses primary documentation and published research where possible. Product capabilities and prices can change; verify vendor details before procurement or production changes.