Skip to content

Test and tune the planner's web search and fetch use - #712

Open
IZO-Ong wants to merge 17 commits into
OpenFn:mainfrom
IZO-Ong:search-test
Open

IZO-Ong wants to merge 17 commits into
OpenFn:mainfrom
IZO-Ong:search-test

Conversation

@IZO-Ong

@IZO-Ong IZO-Ong commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Short Description

Adds a test layer that records how the global chat planner uses web_search / web_fetch, uses it to run a staged experiment on the planner's prompt and config, and ships a prompt line telling the planner to search before fetching.

Fixes #694

Implementation Details

What ships to production

One bullet appended to planner_web_tools_prompt in services/global_chat/prompts.yaml:

web_fetch can only open a URL that already appears in this conversation: in the user's message, or in an earlier search or fetch result. URLs you remember, and URLs in these instructions, are refused. Search first, then fetch a URL from the results.

max_content_tokens, max_uses, the allowlist and planner code are unchanged.

Test layer (services/global_chat/tests/web_tools/)

The questions in #694 are about how the tools are used (order, URLs, refusals, re-fetches). As such, the checks assert on a recorded tool-call trace:

  • trace.py: turns the planner's raw API responses into an ordered list of web calls, pairing calls to results by tool_use_id across rounds.
  • recording.py: RecordingPlanner keeps each raw response, and variants are applied to config/prompts in memory.
  • metrics.py: per-run metrics and the stage winner rule.
  • scenarios.py: six scenarios, each with a trap (two controls that should never touch the web, a shallow and a deep FHIR question, an off-allowlist DHIS2 question, a 3-turn conversation).
  • run_web_experiments.py: the staged runner. Runs are cached on disk keyed by a hash of the resolved model, planner config, prompts and scenario.
  • calibrate.py: fetched the FHIR pages at 10k/25k/50k tokens to find where facts land, which is how the scenario facts were chosen.

Experiment results

The test suite does N=3 runs per cell against claude-opus-4-8. "Grounded" means the fact is in the answer and in a fetched page, suggesting it was read and not recalled.

Stage 0 (baseline): every FHIR first turn (9/9) opened by fetching patient.html from memory and being refused (url_not_in_prior_context).

Stage 0 table
Scenario Web calls Refused fetches Grounded Follow-up fetches
control (OpenFn concept) 0 0 – –
control (code edit) 0 0 – –
fhir_shallow 4.3 1.3 0.33 –
fhir_deep 3.3 1.0 0.67 –
off_allowlist (DHIS2) 0.3 0 – –
multi_turn (3 turns) 6.7 2.7 – 1.7

Stage 1 (fetch strategy): in almost every run, the model's first move was to fetch a URL it remembered (patient.html), which web_fetch refuses. With the search-first line (1a), the model searched, fetched the right page and answered from it every time (6/6), against 3/6 for base.

Stage 1 table (FHIR scenarios)
Variant Refused / run Grounded (shallow / deep) Input tokens Seconds
base 1.17 0.33 / 0.67 106k 41
1a: search-first line 1.50 1.0 / 1.0 115k 46
1b: inject entry URLs 2.00 0.67 / 0.67 132k 53
1c: both 1.33 1.0 / 0.33 130k 47
1d: 1a minus the prompt's example URL 1.83 0.67 / 1.0 131k 57

Stage 2 (max_content_tokens, carrying 1a): keep 10k. fhir_deep never fetched the long Patient page at any size (0/9). The model is smart enough to find the short value-set page, so the model routes around truncation when a short page has the answer.

Stage 2 table
Size Grounded Input tokens Seconds
10k (current) 1.0 115k 46
25k 0.83 125k 50
50k 1.0 153k 66

Stage 3 (findings across turns, carrying 1a): with 1a, no follow-up turn re-fetched, and 2/3 multi-turn runs opened with no refused fetch at all. The findings line only added costs.

Stage 3 table
Variant Refused / run Follow-up fetches Input tokens Seconds
base (Stage 0) 2.67 1.67 262k 123
1a 0.33 0 172k 99
1a + "state your findings" 1.33 0 295k 162

Answers to #694:

Question Answer
Are the tools used only when needed? Yes. Controls made 0 web calls in every variant (30/30 runs).
Does it respect the allowlist? Yes. 0 attempts at banned URLs.
url_not_in_prior_context The first refusal comes from training memory and no prompt change removes it in single-turn questions. The search-first line makes recovery reliable (grounded 6/6) and removed it in 2/3 multi-turn runs. URL injection works as a route but cost more.
max_content_tokens Keep 10k. Bigger fetches cost more and had no measurable gains.
Re-fetching across turns Fixed by the addition of the search-first line.

Test results on the shipped prompt

  • Unit: 609 passed, 1 skipped.
  • Strict suite: 6 passed.
  • Acceptance (judged): 3/4 pass.
    • The one failure, fhir_shallow, gets the facts right; the judge marked it down for how the answer reads. Before searching and fetching, the model writes short notes to itself, like "I need to search first", and the planner glues all of them onto the start of the answer with no spaces. The user sees: "…I need to search first.Now I can fetch…"
    • This already happened before this PR (an old answer began "Now I can fetch the R4 page.Based on…"). The new line makes it more noticeable, because the notes now describe the search-first rule. The fix belongs in the planner, so it's in follow-up 1.

Deviations from the design

The strict suite allows up to 2 refused fetches on a first turn, where the design asked for 0. The model nearly always opens a FHIR question by trying a URL it remembers, which gets refused. No prompt change was able to stop this, although the new shipped line saw a smaller amount of 1–2 refusals per run, while the old prompt had runs with 4. In follow-up turns, any refused fetch still fails.

Follow-ups

  1. Planner narration leaks into answers. Perhaps keep only the final round's text in response/history, or separate rounds with a delimiter.
  2. The same page is often fetched 2–3 times within one turn at every content size.
  3. Answers built from search snippets can't be checked by the grounding metric as search results come back encrypted.

Steps to run test suites

The live suites need an ANTHROPIC_API_KEY for api.anthropic.com.

export ANTHROPIC_API_KEY=sk-ant-...

poetry run pytest services/global_chat/tests/integration/test_web_tools_pass_fail.py -v   # strict suite, ~$1.50
poetry run pytest services/global_chat/tests/acceptance/web_tools -v -s                    # judged specs, ~$2

# re-run an experiment stage (free when cached; --scenario X --runs 1 for one scenario)
cd services && poetry run python -m global_chat.tests.web_tools.run_web_experiments --stage 1 --carry base

AI Usage

Please disclose whether you've used AI in this work (it's cool, we just want to
know!):

  • Yes, I have used AI
  • No, I have not used AI

You can read more details in our
Responsible AI Policy

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Web search: Create tests and optimise tool use behaviour

1 participant