Repository navigation
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Short Description
Adds a test layer that records how the global chat planner uses
web_search/web_fetch, uses it to run a staged experiment on the planner's prompt and config, and ships a prompt line telling the planner to search before fetching.Fixes #694
Implementation Details
What ships to production
One bullet appended to
planner_web_tools_promptinservices/global_chat/prompts.yaml:max_content_tokens,max_uses, the allowlist and planner code are unchanged.Test layer (
services/global_chat/tests/web_tools/)The questions in #694 are about how the tools are used (order, URLs, refusals, re-fetches). As such, the checks assert on a recorded tool-call trace:
trace.py: turns the planner's raw API responses into an ordered list of web calls, pairing calls to results bytool_use_idacross rounds.recording.py:RecordingPlannerkeeps each raw response, and variants are applied to config/prompts in memory.metrics.py: per-run metrics and the stage winner rule.scenarios.py: six scenarios, each with a trap (two controls that should never touch the web, a shallow and a deep FHIR question, an off-allowlist DHIS2 question, a 3-turn conversation).run_web_experiments.py: the staged runner. Runs are cached on disk keyed by a hash of the resolved model, planner config, prompts and scenario.calibrate.py: fetched the FHIR pages at 10k/25k/50k tokens to find where facts land, which is how the scenario facts were chosen.Experiment results
The test suite does N=3 runs per cell against
claude-opus-4-8. "Grounded" means the fact is in the answer and in a fetched page, suggesting it was read and not recalled.Stage 0 (baseline): every FHIR first turn (9/9) opened by fetching
patient.htmlfrom memory and being refused (url_not_in_prior_context).Stage 0 table
Stage 1 (fetch strategy): in almost every run, the model's first move was to fetch a URL it remembered (patient.html), which web_fetch refuses. With the search-first line (1a), the model searched, fetched the right page and answered from it every time (6/6), against 3/6 for base.
Stage 1 table (FHIR scenarios)
Stage 2 (
max_content_tokens, carrying 1a): keep 10k.fhir_deepnever fetched the long Patient page at any size (0/9). The model is smart enough to find the short value-set page, so the model routes around truncation when a short page has the answer.Stage 2 table
Stage 3 (findings across turns, carrying 1a): with 1a, no follow-up turn re-fetched, and 2/3 multi-turn runs opened with no refused fetch at all. The findings line only added costs.
Stage 3 table
Answers to #694:
url_not_in_prior_contextmax_content_tokensTest results on the shipped prompt
Deviations from the design
The strict suite allows up to 2 refused fetches on a first turn, where the design asked for 0. The model nearly always opens a FHIR question by trying a URL it remembers, which gets refused. No prompt change was able to stop this, although the new shipped line saw a smaller amount of 1–2 refusals per run, while the old prompt had runs with 4. In follow-up turns, any refused fetch still fails.
Follow-ups
response/history, or separate rounds with a delimiter.Steps to run test suites
The live suites need an
ANTHROPIC_API_KEYforapi.anthropic.com.AI Usage
Please disclose whether you've used AI in this work (it's cool, we just want to
know!):
You can read more details in our
Responsible AI Policy