Reasoning tokens are output you never see

OpenAI's help article on tokens states that reasoning tokens are not visible as answer text but count toward output usage and are billed as output tokens. Output is the expensive side of the price list: on 2026-10-09, gpt-6.1-sol is $2.00 per million input tokens and $10.00 per million output tokens at standard short-context rates.

OpenAI's reasoning guide shows a sample usage object with 75 input tokens and 1,186 output tokens, of which 1,024 are reasoning tokens. The visible answer is 162 tokens. At $2.00 input and $10.00 output per million, a cost estimate from the visible text is about $1.77 per 1,000 requests, while the billed usage is about $12.01, close to seven times higher.

The same guide warns that a request can stop with status incomplete when it hits max_output_tokens, possibly before any visible output, so you can pay for input and reasoning tokens and receive no answer. OpenAI recommends reserving at least 25,000 tokens for reasoning and output when you start experimenting with these models.

Conversations re-bill every earlier turn

The model does not remember earlier turns for free. With Chat Completions you resend the message history on every request. With the Responses API and previous_response_id, OpenAI's conversation state guide says all previous input tokens for responses in the chain are billed as input tokens.

Take a 10-turn chat with a 1,000-token system prompt, 200-token user messages, and 300-token replies. The transcript the user sees is 5,000 tokens. Turn k sends the system prompt, k user messages, and k minus 1 earlier replies, so the 10 requests bill 34,500 input tokens plus 3,000 output tokens. On gpt-6.1-sol at standard uncached rates that is about $0.099 for the conversation, against about $0.036 if you priced only the system prompt, the user messages, and the replies once.

Prompt caching narrows the gap because the repeated prefix is billed at the cached rate, $0.10 per million on gpt-6.1-sol instead of $2.00. The first write of a cached prefix is priced at the cache-write rate, $2.50 per million on the same model.

Input you did not write: tools, formatting, images, extra completions

A plain-text token count does not include everything in a request. OpenAI's help article says message structure, tools, schemas, images, and files can affect the full input count, and that its input-token counting API includes formatting tokens for message roles and boundaries. Function definitions and MCP server tool lists are sent with every request that includes them, so a large tool catalog is a fixed input cost on each call.

Some parameters multiply output. For Chat Completions, setting n above 1 generates several choices and you are charged for the tokens in all of them. In the legacy Completions API, best_of = 3 can generate up to three times max_tokens across candidates that are not all returned.

  • Tool and function schemas: billed as input on every request that carries them.
  • Message formatting: role and boundary tokens that local tokenizers do not count.
  • Images and files: tokenized by size and detail level, not by characters.
  • n greater than 1: every generated choice is billed, returned to the user or not.

Charges that sit outside the token columns

Even a correct token count misses fees that the pricing page lists separately. Web search is $10.00 per 1,000 calls, with search content tokens billed at model rates on top for most models. Requests above the short-context tier move every token to the long-context column, which for gpt-6.1-sol is $4.00 input and $15.00 output. Fast mode bills gpt-6.1-sol at $4.00 and $20.00, twice standard, and Ultrafast at $12.00 and $60.00. Regional processing (data residency) and FedRAMP endpoints add a 10% uplift for models released on or after March 5, 2026.

To reconcile a bill, read the usage object on each response rather than estimating from text: output_tokens_details.reasoning_tokens shows hidden reasoning and input_tokens_details.cached_tokens shows what the cache covered. Log those fields with the model, endpoint, and service tier, then compare daily totals with the Usage Dashboard. FrugalAI records this per request across providers and can enforce a budget per key, but the same reconciliation works from your own logs.

Frequently asked questions

Are OpenAI reasoning tokens billed?

Yes. OpenAI's documentation says reasoning tokens are not visible through the API but occupy context window space and are billed as output tokens. They appear in the usage object under output_tokens_details.reasoning_tokens. Setting max_output_tokens caps reasoning plus visible output.

Why does a long chat cost more per message over time?

Each request resends or references the earlier turns, and OpenAI bills those earlier input tokens again, including when you chain with previous_response_id. A 10-turn chat with a 1,000-token system prompt, 200-token messages, and 300-token replies bills 34,500 input tokens, not 3,000.

How do I count the real input tokens of an OpenAI request?

Use OpenAI's input-token counting API, which accepts the same input as the Responses API, including tools, images, files, and conversations, and includes formatting tokens. tiktoken counts plain text only and undercounts structured requests.

Sources and further reading

  1. OpenAI Help: Understanding and counting tokens (verified 2026-10-09)
  2. OpenAI reasoning models guide (verified 2026-10-09)
  3. OpenAI conversation state guide (verified 2026-10-09)
  4. OpenAI API pricing (verified 2026-10-09)

FrugalAI uses primary documentation and published research where possible. Product capabilities and prices can change; verify vendor details before procurement or production changes.