Skip to content

Map OpenAI cache write tokens into usage details - #207

Draft
flore2003 wants to merge 1 commit into
mainfrom
cursor/openai-cache-write-tokens-890c
Draft

flore2003 wants to merge 1 commit into
mainfrom
cursor/openai-cache-write-tokens-890c

Conversation

@flore2003

Copy link
Copy Markdown
Member

What

@core-ai/openai now reads prompt-cache writes from the API usage payload:

  • Responses: usage.input_tokens_details.cache_write_tokens
  • Chat Completions, including the streaming usage chunk: usage.prompt_tokens_details.cache_write_tokens

Those values map to usage.inputTokenDetails.cacheWriteTokens. A missing field still maps to 0. inputTokens stays the provider total, inclusive of cache reads and writes.

@core-ai/azure-openai uses these same adapters, so GPT-5.6 and later cache-write billing is reported there as well. The ChatUsage comment that said only Anthropic reports cache writes is updated to match.

Why

Starting with GPT-5.6, OpenAI bills cache writes at 1.25× the uncached input rate and reports them in cache_write_tokens. The mappers hardcoded cacheWriteTokens: 0, so the written tokens stayed inside inputTokens with no write breakdown. Callers that price cacheWriteTokens separately under-reported cost by 0.25× input on every written token.

Tests

Adapter fixtures cover Responses and Chat Completions, streaming and non-streaming, for both a present cache_write_tokens value and a missing field. npm run test passed for @core-ai/openai, @core-ai/core-ai, and @core-ai/azure-openai. Lint and tsc --noEmit passed for the OpenAI and core packages.

Open in Web Open in Cursor 

GPT-5.6 and later report prompt-cache writes in cache_write_tokens.
Read that field for the Responses API and Chat Completions, including
streaming usage chunks, so cacheWriteTokens is no longer hardcoded to 0.

Co-authored-by: florian <florian@omnifact.ai>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants