Skip to content

Development plan: recording, the one skill install path, and the headless validation that motivated them #78

Description

@aarontuor

Tracking issue for the work that came out of validating the walkthroughs headless. Each section below is a pull request or a short series of them, open or planned. The branch comprehensive-review holds every planned change, one concern per commit, so any section can be read before its pull request opens.

Terms

  • Walkthrough: a tutorial under use_cases/. Its README has a list of prompts a user types to an agent, and a list of post-conditions that say what the project should hold afterwards.
  • Headless run: a script sends the walkthrough's prompts to Claude Code or Codex with no person present, and the result is judged against the post-conditions.
  • Record: the file dsagt-run writes when the agent runs a command through it: the command, when it ran, whether it succeeded, what it printed, which files it read and wrote.
  • Reconstructed pipeline script: the shell script dsagt builds from the records (reconstruct_pipeline, often saved as pipeline.sh), which repeats the work on a fresh copy of the project.
  • Registered command: a reusable program the agent has saved a description of (name, purpose, parameters), so it can be found and run again.
  • Skill: a folder with a SKILL.md of instructions for the agent and, often, a scripts/ folder of programs the instructions tell it to run.

Source of the work

The seven walkthroughs were run headless with Claude Code (sonnet-4-5) and Codex (gpt-5.5) on 2026-09-16 and 2026-09-17. Each run was judged by the README's post-conditions, by the records, and by whether the reconstructed pipeline script reproduced the outputs. Two opposite failures came up in most walkthroughs:

  1. Work that should have been recorded was not. Claude ran a skill's programs directly, as the skill's text showed them, computed numbers it reported in inline python, and wrote one-off scripts outside the project. One walkthrough finished with every output file present and no records.
  2. Work that needed no program got one. Codex followed the rule "every operation that writes a data file runs as a registered command" to the letter: it wrote and registered a 403-line program to produce one document, and ran the data-quality check four times on one unchanged 8-row file.

Each rewording of the instructions fixed one agent and made the other worse. The sections below change how dsagt works so the rule the agent follows is short: a command that produces or transforms a dataset file runs as a registered code, and a skill's programs are registered the moment the skill is in the project.

Post-conditions met by Claude, before and after the changes:

Walkthrough 09-17 (main + #72, #74, #75) 09-18 (comprehensive-review)
aidrin-ai-readiness 6/6 5/6 (the post-condition now requires the datacard to pass its validator)
genesis_skills 3/5, 0 records 6/6, 9 records
vasp_dft 6/8, 8 unrecorded python runs 7/8, 0 unrecorded python runs
tokamak_stability 4 + 2 partly of 6; the plots and the reported safety factor came from unrecorded scripts 5 + 1 partly of 6; both came from registered commands
microbial_isolates 3 + 4 partly of 7 3 + 2 partly of 7 (a headless turn ends a command Claude Code has moved to the background at ten minutes; an interactive session lets it finish)
cryoem 7/8 6/8; codex 7/8 through prompt 9, cut by the usage limit (the datacard is the miss; see Upstream)
combustion_simulation 7/7 6 + 1 partly of 7; the agent wrote its converter before seeing the reference output, then corrected it against the reference to an exact match, which is the walkthrough's intent

Reconstructed pipeline scripts, run on a fresh copy: on 09-17 four of seven failed (missing report files, missing folders, a program that refuses to overwrite). On 09-18 six of seven exit 0. Logged conversation turns for the 10 prompts of one walkthrough: 54 on 09-17, 9 on 09-18 (section 9).

Sections

1. Headless walkthrough validation

Problem. The walkthroughs were checked by hand. A change to the agent's instructions, to a tool's description, or to a built-in skill changes what agents do, and the unit tests cannot show that.

Change.

2. Complete records, a narrower rule, and a faster dsagt-run

Background. dsagt keeps a log of the pipeline steps an agent runs on the user's data. A registered code's stored command starts with dsagt-run, which runs the command and writes a record: the command, when it ran, whether it succeeded, what it printed, and which files it read and wrote. dsagt turns the records into a shell script that repeats the work. The agent's other commands, its inspections and comparisons, are already kept with their output in the conversation traces.

Problem. Claude ran a skill's programs directly, as the skill's text showed them, so pipeline steps left no record (section 4 fixes the text). Codex read the rule "every operation that writes a data artifact or an audit report runs as a registered code" to include a document, and wrote and registered a 403-line program to produce one datacard. The records that did exist were not complete enough to replay: printed reports were saved outside the record, records of codes without declared file parameters named no files, a run ended from outside left no record, and four of seven reconstructed pipeline scripts failed on a fresh copy.

Change. Planned, one pull request with section 3.

  • The rule becomes: a command that produces or transforms a dataset file is a pipeline step, and it runs as a registered code by its stored line. A document is not a dataset file.
  • Printed reports. The record holds what the command printed, in full, so a report a code prints is read back from its record.
  • Killed commands. A command ended from outside, as when a time limit ends the agent's turn, leaves a record that says so.
  • File lists. A record lists the files the command read and wrote, and the reconstructed script is ordered by them. When a code declares no input and output parameters, or the agent spelled a flag differently from the saved description (--in1 for -i), or left off the uv run prefix, dsagt infers the files from the command line: a named file or directory that exists before the run is an input; one that exists only afterwards, or that the run changed, is an output. A declared side that comes up empty is filled from the command line on its own, rather than the whole inference being switched off.
  • Loop scripts. When a registered loop script itself runs registered codes, the reconstructed pipeline script runs the inner commands only through the outer one, unless the outer one was killed.
  • The reconstructed pipeline script creates its folders first, deletes a file before the step that rewrites it, and can be saved with reconstruct_pipeline(output=...).
  • Speed. dsagt-run added 1.3 to 2 seconds to every command. It adds about 0.2 seconds: the trace is logged by a separate process after the command ends.
  • One option. dsagt-run --code <name> -- <command>, and nothing else. --input-files, --output-files, --session, --record-id and --records-dir are gone; each value is derived, and no agent command in 1248 records passed any of them.

Tried and withdrawn. For a day the branch also recorded any command with no registration (dsagt-run -- <command>), under a rule that every computation leaves a record, and kept a copy of one-off scripts so they would replay. In 21 runs, 78 of the 81 records this produced were inspection one-liners that wrote no file, and they made up most of a reconstructed script's steps. How the reconstruction should cover a one-off shell step the pipeline depends on (an unpack, a move) stays open.

3. The data-quality check runs on tables, once per file

Background. dsagt asks the agent to run AIDRIN, a data-quality tool for tables, before and after each step that changes a table, and to keep the reports.

Problem.

  • The records of the AIDRIN runs named no files, and the report was saved by a shell redirect that the record did not include, so neither the record nor the reconstructed pipeline script showed which file was checked or where the report went.
  • The agent repeated the check on files that had not changed.
  • The agent ran it on files that are not tables. AIDRIN scored a simulation output file (HDF5, many arrays) as a 20-column table without complaint. The AIDRIN side of this is reported upstream.

Change. Planned, one pull request (1bf4eb2), after section 2.

  • Every record holds a checksum of each file the command read, taken before the run, and of each file it wrote, taken after.
  • A new tool, readiness_reports(path), returns the AIDRIN report on record for a file: the one made while the file had its present content, with the text the check printed, and the reports from before the file changed.
  • The instructions say what counts as a table (CSV, TSV, Excel, JSON, HDF5, Parquet, npz), warn that a JSON, HDF5 or NumPy file may hold nested structure the baseline reads as one flat table, say the report is saved to the run's record, and say to run the check on the data as it arrives and on the output of each transformation. A score compares with an earlier report on the same table, or with the report of the table it was derived from.
  • A test keeps the copy of this paragraph on the documentation site equal to the one the agent reads.

4. One way for a skill's programs to become registered commands

Problem. A skill reaches a project four ways: dsagt installs three built-in skills when the project is created; the agent installs one from a catalog; the agent writes its own; or the agent registers a single program. Each way had its own code, and they did different things. The built-in skills' programs were registered from hand-written entries. A catalog skill's programs, and the agent's own, were registered only if the agent thought to do it. The skill's text kept telling the agent to run the programs directly (python3 scripts/introspect.py), which leaves no record, and Claude did what the text said.

Change. Planned, one pull request.

  • scan-directory, the one example command packaged inside dsagt, is removed with the code that supported packaged commands (8980171). It listed files, and taught the agent that reading the filesystem needs a registered command. The built-in skills' programs are the examples now.
  • One function, skills.register_skill_scripts, serves all four ways (c626ca7, 53a12f3, part of ac257fa). It registers every runnable program under the skill's scripts/, taking the parameters from a hand-written entry if there is one and otherwise from the program's own argument definitions. It rewrites the direct invocations in the skill's text to the recorded form, and replies to the agent with the lines to use. Registered commands and skills now share one folder, skills/: a registered command is a skill whose header names an executable. A test holds all four ways to the same result.
  • dsagt gives each agent platform the project's skills in the folder that platform reads (.claude/skills, .agents/skills), as a symbolic link per skill (c2c7fd4). As a real copy, a skill existed twice and the two drifted apart the moment either changed: under Claude the platform's copy went stale, and Codex kept them in step by editing both by hand, 14 times in one run. With a link there is one copy, so an edit to a skill, by the agent or by a person, is what every reader sees.
  • save_code_spec under a name a script already has replies with that name instead of registering the script twice, and under an installed skill's name it is refused, since a code and a skill share one directory. The reply says to save the same spec under the existing name, which is how an argparse-derived spec gains its file roles and dependencies.
  • search_skills describes the catalogs it searches. It said it spanned installed skills, which every agent discovers natively and dsagt does not index.

Overlap with the dataset-builder work (#43). That work places packaged commands in src/dsagt/codes/, which this section removes, and #69 adds its checks as a skill whose program imports from the dsagt package. The rule now in CLAUDE.md is: tools on the dsagt server own dsagt's own state; a packaged program that works on the user's files runs through dsagt-run and imports nothing from dsagt. Under that rule the staleness check is a server tool and the rest is a skill registered through this section's function. Branch 46-check-dataset-staleness-tool shows that shape as a proposal. The order of #52, #53, #69 against this section is the one open coordination question.

5. Withdrawn: a check at the point where the agent issues a command

A Claude Code hook that put a bare python call under dsagt-run was built and validated, then removed. With section 4 in place the pipeline's steps are recorded without it; what it caught was inspection one-liners, which the conversation traces already hold. It covered one agent platform of five, edited the user's .claude/settings.json, and changed a command without the agent knowing. Codex, with no hook, recorded its pipeline steps from the instructions alone.

6. The agent can find past runs; three duplicate tools go

Background. dsagt keeps several searchable collections for the agent: documents the user adds, registered commands, records of past runs, and memories.

Problem. The tool that lists collections left out dsagt's own, so the agent had no way to learn that past runs were searchable, and the search tool could not narrow a search to one command. Separately, the dsagt server offered run_command, read_file and http_request, which every agent platform already has. A command run through the server's run_command runs in the server's environment, not the user's, and leaves no record.

Change. Planned, one pull request (9855d1c). kb_list_collections lists every collection with what it is for, the fields a search can filter on, and its size. kb_search takes a filter. The three tools are removed.

install_dependencies goes with them, and so does the install save_code_spec ran for a spec with dependencies: both installed into the dsagt server's own interpreter, which is not the environment the agent's shell runs a command in, and a spec's dependencies already reach the run through the uv run --with prefix in its stored line. The tool was called in none of 325 agent transcripts. The server has 17 tools.

7. The server sees the user's Python environment under Codex and Cline

Problem. Codex and Cline start the dsagt server with only the environment variables written in their config file, not the ones from the shell the user launched from. The server then ran with a Python that lacked the packages in the user's activated environment.

Change. Planned, one pull request (63a6c88). dsagt init and dsagt start copy a fixed list of variables from the launching shell into that config (PATH, VIRTUAL_ENV, CONDA_PREFIX, PYTHONPATH, the library paths), plus any names the user lists under mcp.env_passthrough. A name that looks like a credential (*_KEY, *_TOKEN, *_SECRET) is refused, and a test holds that.

The embedding backend's own credential follows the same rule. It read EMBEDDING_API_KEY, then LLM_API_KEY, then OPENAI_API_KEY, and its base URL from EMBEDDING_BASE_URL or OPENAI_BASE_URL; the server also read embedding.api_key from the project config. The OPENAI_ names are the agent's LLM-provider credentials, which dsagt never reads. One name each now, from the shell or ~/.config/dsagt/env.

8. dsagt init works without git

Problem. dsagt init downloads the built-in skills from GitHub and needed the git program to do it. A user who installs dsagt with pip may not have git.

Change. Planned, one pull request (579a8d2). The skills are downloaded as an archive over HTTPS, and the exact commit is still written down beside them. git is used only as a fallback, for a private repository that the download is refused for.

9. A resumed conversation is logged once

Background. dsagt reads the agent's conversation file from disk and logs each turn (prompt, tool calls, tokens) to a local MLflow store, so dsagt info can show what a session did and cost.

Problem. Each claude --continue or codex exec resume starts a new dsagt session on the same conversation file, and dsagt logged every earlier turn again under the new session. One 10-prompt walkthrough showed 54 turns, and dsagt info counted the tokens several times over.

Change. Open: #77. dsagt remembers which turns it has logged per conversation file, not per session, and leaves a turn that is still in progress for the next pass. Open: #76, a one-line fix to an integration test, found when the full test suite was run before the validation.

10. Documents and comments brought in line with the code

Problem. The README, the documentation site, and the comments in the code had drifted from what the code does, and from the project's writing rules.

Change. Planned, last. 18 false statements corrected, one test that checked nothing removed, the writing rules applied across the source, tests, site and walkthroughs, and the changelog entries for everything above.

A second pass on 2026-09-21 covered what the branch had since made false: the module docstrings for observability.py, registry.py and provenance.py; the agent card, which advertised an embedder, a reranker, a tool count and a tool argument the code does not have; the per-operation check rule, which still told the agent to write reports to audit/ directly above the paragraph saying the report is the run's record; and the use-case catalog, which the site generates. It also removed two hand-tests describing an agent dsagt does not support, an API key committed to tests/test_site_config.yaml, and a failing integration test whose integration marker had kept it out of the green run.

The pass lists, for a later one and not changed here: four pieces of code nothing calls, nine alive only in their tests, eleven places where an error is caught and dropped, dsagt start --agent which is declared and never read, and dsagt rm --all which is undocumented and untested.

Upstream

Order

Sections 1 (first part) and 9 are open now as #74, #76 and #77. The rest is five pull requests, stacked because they change many of the same files (provenance.py, skills.py, registry_tools.py, the agent's instructions), each based on the one before:

  1. dsagt-run and the record: sections 2 and 3, and the wrapper's one option.
  2. Server and init: sections 6, 7 and 8, the project-registry lock and the small trace fixes.
  3. Skills and codes: section 4, with the install-path fixes found by validating the branch.
  4. Walkthroughs: the text fixes, one download per walkthrough, the headless-runs guide, and a checker that scores a headless run's dsagt checks apart from the agent's outcomes (section 1, second part).
  5. Documents and changelog: section 10.

Each description ties its problem and fix to its commits and points back here for the background. The branch comprehensive-review holds the same tree as the top of the stack.

How the walkthroughs are judged changed on 2026-09-19. Three runs of each walkthrough on one tree showed a post-condition count moving by one or two between identical runs, and a third of the misses came from the headless harness itself (a turn ending a backgrounded command, nobody to answer a skill's second question). Mechanical checks of dsagt now fail a run; agent-dependent outcomes are reported as rates over three or more runs; what a headless session cannot show is not scored. The before-and-after table above predates that and is kept as the record of what started the work.

Tabled

Under consideration, not planned: an interface for a person, whose controls are the same tools the agent holds.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions