Skip to content

Cache planner history across tool-loop rounds - #713

Open
IZO-Ong wants to merge 2 commits into
OpenFn:mainfrom
IZO-Ong:web-search-cache
Open

IZO-Ong wants to merge 2 commits into
OpenFn:mainfrom
IZO-Ong:web-search-cache

Conversation

@IZO-Ong

@IZO-Ong IZO-Ong commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Short Description

Adds a top-level automatic prompt caching to each planner API call,. This allows later rounds of the tool loop to read earlier rounds, including web search and fetch results, from cache.

Fixes #693

Implementation Details

The original planner only cached the tool definitions and system prompt. The API already caches web results inside a single call, but each new round of the planner's tool loop (tool_use rounds, pause_turn continuations, the no-tools wrap-up round, the web-downgrade retry) re-sent the whole messages array with no breakpoint in it, so fetched pages were billed again on every round.

Changes:

  • services/global_chat/planner.py: a _HISTORY_CACHE constant ({"cache_control": {"type": "ephemeral"}}) passed into both the streaming and non-streaming calls in _call_api. The API puts the breakpoint on the last cacheable block and moves it forward each round. This is applied to every planner call.
  • Each request now uses 3 of the API's 4.

Tests:

  • Four new unit tests in test_planner.py check that cache_control is sent on streamed and non-streamed calls, on every round of a pause -> tool_use -> wrap-up turn, and on the web-downgrade retry.

AI Usage

Please disclose whether you've used AI in this work (it's cool, we just want to
know!):

  • Yes, I have used AI
  • No, I have not used AI

You can read more details in our
Responsible AI Policy

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Web search: Extend cache control to search results

1 participant