Tracking issue for the work that came out of validating the walkthroughs headless. Each section below is a pull request or a short series of them, open or planned. The branch comprehensive-review holds every planned change, one concern per commit, so any section can be read before its pull request opens.
Terms
- Walkthrough: a tutorial under
use_cases/. Its README has a list of prompts a user types to an agent, and a list of post-conditions that say what the project should hold afterwards.
- Headless run: a script sends the walkthrough's prompts to Claude Code or Codex with no person present, and the result is judged against the post-conditions.
- Record: the file
dsagt-run writes when the agent runs a command through it: the command, when it ran, whether it succeeded, what it printed, which files it read and wrote.
- Reconstructed pipeline script: the shell script dsagt builds from the records (
reconstruct_pipeline, often saved as pipeline.sh), which repeats the work on a fresh copy of the project.
- Registered command: a reusable program the agent has saved a description of (name, purpose, parameters), so it can be found and run again.
- Skill: a folder with a
SKILL.md of instructions for the agent and, often, a scripts/ folder of programs the instructions tell it to run.
Source of the work
The seven walkthroughs were run headless with Claude Code (sonnet-4-5) and Codex (gpt-5.5) on 2026-09-16 and 2026-09-17. Each run was judged by the README's post-conditions, by the records, and by whether the reconstructed pipeline script reproduced the outputs. Two opposite failures came up in most walkthroughs:
- Work that should have been recorded was not. Claude ran a skill's programs directly, as the skill's text showed them, computed numbers it reported in inline python, and wrote one-off scripts outside the project. One walkthrough finished with every output file present and no records.
- Work that needed no program got one. Codex followed the rule "every operation that writes a data file runs as a registered command" to the letter: it wrote and registered a 403-line program to produce one document, and ran the data-quality check four times on one unchanged 8-row file.
Each rewording of the instructions fixed one agent and made the other worse. The sections below change how dsagt works so the rule the agent follows is short: a command that produces or transforms a dataset file runs as a registered code, and a skill's programs are registered the moment the skill is in the project.
Post-conditions met by Claude, before and after the changes:
| Walkthrough |
09-17 (main + #72, #74, #75) |
09-18 (comprehensive-review) |
| aidrin-ai-readiness |
6/6 |
5/6 (the post-condition now requires the datacard to pass its validator) |
| genesis_skills |
3/5, 0 records |
6/6, 9 records |
| vasp_dft |
6/8, 8 unrecorded python runs |
7/8, 0 unrecorded python runs |
| tokamak_stability |
4 + 2 partly of 6; the plots and the reported safety factor came from unrecorded scripts |
5 + 1 partly of 6; both came from registered commands |
| microbial_isolates |
3 + 4 partly of 7 |
3 + 2 partly of 7 (a headless turn ends a command Claude Code has moved to the background at ten minutes; an interactive session lets it finish) |
| cryoem |
7/8 |
6/8; codex 7/8 through prompt 9, cut by the usage limit (the datacard is the miss; see Upstream) |
| combustion_simulation |
7/7 |
6 + 1 partly of 7; the agent wrote its converter before seeing the reference output, then corrected it against the reference to an exact match, which is the walkthrough's intent |
Reconstructed pipeline scripts, run on a fresh copy: on 09-17 four of seven failed (missing report files, missing folders, a program that refuses to overwrite). On 09-18 six of seven exit 0. Logged conversation turns for the 10 prompts of one walkthrough: 54 on 09-17, 9 on 09-18 (section 9).
Sections
1. Headless walkthrough validation
Problem. The walkthroughs were checked by hand. A change to the agent's instructions, to a tool's description, or to a built-in skill changes what agents do, and the unit tests cannot show that.
Change.
2. Complete records, a narrower rule, and a faster dsagt-run
Background. dsagt keeps a log of the pipeline steps an agent runs on the user's data. A registered code's stored command starts with dsagt-run, which runs the command and writes a record: the command, when it ran, whether it succeeded, what it printed, and which files it read and wrote. dsagt turns the records into a shell script that repeats the work. The agent's other commands, its inspections and comparisons, are already kept with their output in the conversation traces.
Problem. Claude ran a skill's programs directly, as the skill's text showed them, so pipeline steps left no record (section 4 fixes the text). Codex read the rule "every operation that writes a data artifact or an audit report runs as a registered code" to include a document, and wrote and registered a 403-line program to produce one datacard. The records that did exist were not complete enough to replay: printed reports were saved outside the record, records of codes without declared file parameters named no files, a run ended from outside left no record, and four of seven reconstructed pipeline scripts failed on a fresh copy.
Change. Planned, one pull request with section 3.
- The rule becomes: a command that produces or transforms a dataset file is a pipeline step, and it runs as a registered code by its stored line. A document is not a dataset file.
- Printed reports. The record holds what the command printed, in full, so a report a code prints is read back from its record.
- Killed commands. A command ended from outside, as when a time limit ends the agent's turn, leaves a record that says so.
- File lists. A record lists the files the command read and wrote, and the reconstructed script is ordered by them. When a code declares no input and output parameters, or the agent spelled a flag differently from the saved description (
--in1 for -i), or left off the uv run prefix, dsagt infers the files from the command line: a named file or directory that exists before the run is an input; one that exists only afterwards, or that the run changed, is an output. A declared side that comes up empty is filled from the command line on its own, rather than the whole inference being switched off.
- Loop scripts. When a registered loop script itself runs registered codes, the reconstructed pipeline script runs the inner commands only through the outer one, unless the outer one was killed.
- The reconstructed pipeline script creates its folders first, deletes a file before the step that rewrites it, and can be saved with
reconstruct_pipeline(output=...).
- Speed.
dsagt-run added 1.3 to 2 seconds to every command. It adds about 0.2 seconds: the trace is logged by a separate process after the command ends.
- One option.
dsagt-run --code <name> -- <command>, and nothing else. --input-files, --output-files, --session, --record-id and --records-dir are gone; each value is derived, and no agent command in 1248 records passed any of them.
Tried and withdrawn. For a day the branch also recorded any command with no registration (dsagt-run -- <command>), under a rule that every computation leaves a record, and kept a copy of one-off scripts so they would replay. In 21 runs, 78 of the 81 records this produced were inspection one-liners that wrote no file, and they made up most of a reconstructed script's steps. How the reconstruction should cover a one-off shell step the pipeline depends on (an unpack, a move) stays open.
3. The data-quality check runs on tables, once per file
Background. dsagt asks the agent to run AIDRIN, a data-quality tool for tables, before and after each step that changes a table, and to keep the reports.
Problem.
- The records of the AIDRIN runs named no files, and the report was saved by a shell redirect that the record did not include, so neither the record nor the reconstructed pipeline script showed which file was checked or where the report went.
- The agent repeated the check on files that had not changed.
- The agent ran it on files that are not tables. AIDRIN scored a simulation output file (HDF5, many arrays) as a 20-column table without complaint. The AIDRIN side of this is reported upstream.
Change. Planned, one pull request (1bf4eb2), after section 2.
- Every record holds a checksum of each file the command read, taken before the run, and of each file it wrote, taken after.
- A new tool,
readiness_reports(path), returns the AIDRIN report on record for a file: the one made while the file had its present content, with the text the check printed, and the reports from before the file changed.
- The instructions say what counts as a table (CSV, TSV, Excel, JSON, HDF5, Parquet, npz), warn that a JSON, HDF5 or NumPy file may hold nested structure the baseline reads as one flat table, say the report is saved to the run's record, and say to run the check on the data as it arrives and on the output of each transformation. A score compares with an earlier report on the same table, or with the report of the table it was derived from.
- A test keeps the copy of this paragraph on the documentation site equal to the one the agent reads.
4. One way for a skill's programs to become registered commands
Problem. A skill reaches a project four ways: dsagt installs three built-in skills when the project is created; the agent installs one from a catalog; the agent writes its own; or the agent registers a single program. Each way had its own code, and they did different things. The built-in skills' programs were registered from hand-written entries. A catalog skill's programs, and the agent's own, were registered only if the agent thought to do it. The skill's text kept telling the agent to run the programs directly (python3 scripts/introspect.py), which leaves no record, and Claude did what the text said.
Change. Planned, one pull request.
scan-directory, the one example command packaged inside dsagt, is removed with the code that supported packaged commands (8980171). It listed files, and taught the agent that reading the filesystem needs a registered command. The built-in skills' programs are the examples now.
- One function,
skills.register_skill_scripts, serves all four ways (c626ca7, 53a12f3, part of ac257fa). It registers every runnable program under the skill's scripts/, taking the parameters from a hand-written entry if there is one and otherwise from the program's own argument definitions. It rewrites the direct invocations in the skill's text to the recorded form, and replies to the agent with the lines to use. Registered commands and skills now share one folder, skills/: a registered command is a skill whose header names an executable. A test holds all four ways to the same result.
- dsagt gives each agent platform the project's skills in the folder that platform reads (
.claude/skills, .agents/skills), as a symbolic link per skill (c2c7fd4). As a real copy, a skill existed twice and the two drifted apart the moment either changed: under Claude the platform's copy went stale, and Codex kept them in step by editing both by hand, 14 times in one run. With a link there is one copy, so an edit to a skill, by the agent or by a person, is what every reader sees.
save_code_spec under a name a script already has replies with that name instead of registering the script twice, and under an installed skill's name it is refused, since a code and a skill share one directory. The reply says to save the same spec under the existing name, which is how an argparse-derived spec gains its file roles and dependencies.
search_skills describes the catalogs it searches. It said it spanned installed skills, which every agent discovers natively and dsagt does not index.
Overlap with the dataset-builder work (#43). That work places packaged commands in src/dsagt/codes/, which this section removes, and #69 adds its checks as a skill whose program imports from the dsagt package. The rule now in CLAUDE.md is: tools on the dsagt server own dsagt's own state; a packaged program that works on the user's files runs through dsagt-run and imports nothing from dsagt. Under that rule the staleness check is a server tool and the rest is a skill registered through this section's function. Branch 46-check-dataset-staleness-tool shows that shape as a proposal. The order of #52, #53, #69 against this section is the one open coordination question.
5. Withdrawn: a check at the point where the agent issues a command
A Claude Code hook that put a bare python call under dsagt-run was built and validated, then removed. With section 4 in place the pipeline's steps are recorded without it; what it caught was inspection one-liners, which the conversation traces already hold. It covered one agent platform of five, edited the user's .claude/settings.json, and changed a command without the agent knowing. Codex, with no hook, recorded its pipeline steps from the instructions alone.
6. The agent can find past runs; three duplicate tools go
Background. dsagt keeps several searchable collections for the agent: documents the user adds, registered commands, records of past runs, and memories.
Problem. The tool that lists collections left out dsagt's own, so the agent had no way to learn that past runs were searchable, and the search tool could not narrow a search to one command. Separately, the dsagt server offered run_command, read_file and http_request, which every agent platform already has. A command run through the server's run_command runs in the server's environment, not the user's, and leaves no record.
Change. Planned, one pull request (9855d1c). kb_list_collections lists every collection with what it is for, the fields a search can filter on, and its size. kb_search takes a filter. The three tools are removed.
install_dependencies goes with them, and so does the install save_code_spec ran for a spec with dependencies: both installed into the dsagt server's own interpreter, which is not the environment the agent's shell runs a command in, and a spec's dependencies already reach the run through the uv run --with prefix in its stored line. The tool was called in none of 325 agent transcripts. The server has 17 tools.
7. The server sees the user's Python environment under Codex and Cline
Problem. Codex and Cline start the dsagt server with only the environment variables written in their config file, not the ones from the shell the user launched from. The server then ran with a Python that lacked the packages in the user's activated environment.
Change. Planned, one pull request (63a6c88). dsagt init and dsagt start copy a fixed list of variables from the launching shell into that config (PATH, VIRTUAL_ENV, CONDA_PREFIX, PYTHONPATH, the library paths), plus any names the user lists under mcp.env_passthrough. A name that looks like a credential (*_KEY, *_TOKEN, *_SECRET) is refused, and a test holds that.
The embedding backend's own credential follows the same rule. It read EMBEDDING_API_KEY, then LLM_API_KEY, then OPENAI_API_KEY, and its base URL from EMBEDDING_BASE_URL or OPENAI_BASE_URL; the server also read embedding.api_key from the project config. The OPENAI_ names are the agent's LLM-provider credentials, which dsagt never reads. One name each now, from the shell or ~/.config/dsagt/env.
8. dsagt init works without git
Problem. dsagt init downloads the built-in skills from GitHub and needed the git program to do it. A user who installs dsagt with pip may not have git.
Change. Planned, one pull request (579a8d2). The skills are downloaded as an archive over HTTPS, and the exact commit is still written down beside them. git is used only as a fallback, for a private repository that the download is refused for.
9. A resumed conversation is logged once
Background. dsagt reads the agent's conversation file from disk and logs each turn (prompt, tool calls, tokens) to a local MLflow store, so dsagt info can show what a session did and cost.
Problem. Each claude --continue or codex exec resume starts a new dsagt session on the same conversation file, and dsagt logged every earlier turn again under the new session. One 10-prompt walkthrough showed 54 turns, and dsagt info counted the tokens several times over.
Change. Open: #77. dsagt remembers which turns it has logged per conversation file, not per session, and leaves a turn that is still in progress for the next pass. Open: #76, a one-line fix to an integration test, found when the full test suite was run before the validation.
10. Documents and comments brought in line with the code
Problem. The README, the documentation site, and the comments in the code had drifted from what the code does, and from the project's writing rules.
Change. Planned, last. 18 false statements corrected, one test that checked nothing removed, the writing rules applied across the source, tests, site and walkthroughs, and the changelog entries for everything above.
A second pass on 2026-09-21 covered what the branch had since made false: the module docstrings for observability.py, registry.py and provenance.py; the agent card, which advertised an embedder, a reranker, a tool count and a tool argument the code does not have; the per-operation check rule, which still told the agent to write reports to audit/ directly above the paragraph saying the report is the run's record; and the use-case catalog, which the site generates. It also removed two hand-tests describing an agent dsagt does not support, an API key committed to tests/test_site_config.yaml, and a failing integration test whose integration marker had kept it out of the green run.
The pass lists, for a later one and not changed here: four pieces of code nothing calls, nine alive only in their tests, eleven places where an error is caught and dropped, dsagt start --agent which is declared and never read, and dsagt rm --all which is undocumented and untested.
Upstream
Order
Sections 1 (first part) and 9 are open now as #74, #76 and #77. The rest is five pull requests, stacked because they change many of the same files (provenance.py, skills.py, registry_tools.py, the agent's instructions), each based on the one before:
dsagt-run and the record: sections 2 and 3, and the wrapper's one option.
- Server and init: sections 6, 7 and 8, the project-registry lock and the small trace fixes.
- Skills and codes: section 4, with the install-path fixes found by validating the branch.
- Walkthroughs: the text fixes, one download per walkthrough, the headless-runs guide, and a checker that scores a headless run's dsagt checks apart from the agent's outcomes (section 1, second part).
- Documents and changelog: section 10.
Each description ties its problem and fix to its commits and points back here for the background. The branch comprehensive-review holds the same tree as the top of the stack.
How the walkthroughs are judged changed on 2026-09-19. Three runs of each walkthrough on one tree showed a post-condition count moving by one or two between identical runs, and a third of the misses came from the headless harness itself (a turn ending a backgrounded command, nobody to answer a skill's second question). Mechanical checks of dsagt now fail a run; agent-dependent outcomes are reported as rates over three or more runs; what a headless session cannot show is not scored. The before-and-after table above predates that and is kept as the record of what started the work.
Tabled
Under consideration, not planned: an interface for a person, whose controls are the same tools the agent holds.
Tracking issue for the work that came out of validating the walkthroughs headless. Each section below is a pull request or a short series of them, open or planned. The branch
comprehensive-reviewholds every planned change, one concern per commit, so any section can be read before its pull request opens.Terms
use_cases/. Its README has a list of prompts a user types to an agent, and a list of post-conditions that say what the project should hold afterwards.dsagt-runwrites when the agent runs a command through it: the command, when it ran, whether it succeeded, what it printed, which files it read and wrote.reconstruct_pipeline, often saved aspipeline.sh), which repeats the work on a fresh copy of the project.SKILL.mdof instructions for the agent and, often, ascripts/folder of programs the instructions tell it to run.Source of the work
The seven walkthroughs were run headless with Claude Code (sonnet-4-5) and Codex (gpt-5.5) on 2026-09-16 and 2026-09-17. Each run was judged by the README's post-conditions, by the records, and by whether the reconstructed pipeline script reproduced the outputs. Two opposite failures came up in most walkthroughs:
Each rewording of the instructions fixed one agent and made the other worse. The sections below change how dsagt works so the rule the agent follows is short: a command that produces or transforms a dataset file runs as a registered code, and a skill's programs are registered the moment the skill is in the project.
Post-conditions met by Claude, before and after the changes:
main+ #72, #74, #75)comprehensive-review)Reconstructed pipeline scripts, run on a fresh copy: on 09-17 four of seven failed (missing report files, missing folders, a program that refuses to overwrite). On 09-18 six of seven exit 0. Logged conversation turns for the 10 prompts of one walkthrough: 54 on 09-17, 9 on 09-18 (section 9).
Sections
1. Headless walkthrough validation
Problem. The walkthroughs were checked by hand. A change to the agent's instructions, to a tool's description, or to a built-in skill changes what agents do, and the unit tests cannot show that.
Change.
tests/headless_usecases.pysends a README's prompts through Claude Code or Codex in one conversation and keeps a log. Open: Headless walkthrough driver; walkthrough fixes from the headless runs; an artifacts review prompt #74.headless-runs-guide).99c527c. Planned, after Headless walkthrough driver; walkthrough fixes from the headless runs; an artifacts review prompt #74.2. Complete records, a narrower rule, and a faster
dsagt-runBackground. dsagt keeps a log of the pipeline steps an agent runs on the user's data. A registered code's stored command starts with
dsagt-run, which runs the command and writes a record: the command, when it ran, whether it succeeded, what it printed, and which files it read and wrote. dsagt turns the records into a shell script that repeats the work. The agent's other commands, its inspections and comparisons, are already kept with their output in the conversation traces.Problem. Claude ran a skill's programs directly, as the skill's text showed them, so pipeline steps left no record (section 4 fixes the text). Codex read the rule "every operation that writes a data artifact or an audit report runs as a registered code" to include a document, and wrote and registered a 403-line program to produce one datacard. The records that did exist were not complete enough to replay: printed reports were saved outside the record, records of codes without declared file parameters named no files, a run ended from outside left no record, and four of seven reconstructed pipeline scripts failed on a fresh copy.
Change. Planned, one pull request with section 3.
--in1for-i), or left off theuv runprefix, dsagt infers the files from the command line: a named file or directory that exists before the run is an input; one that exists only afterwards, or that the run changed, is an output. A declared side that comes up empty is filled from the command line on its own, rather than the whole inference being switched off.reconstruct_pipeline(output=...).dsagt-runadded 1.3 to 2 seconds to every command. It adds about 0.2 seconds: the trace is logged by a separate process after the command ends.dsagt-run --code <name> -- <command>, and nothing else.--input-files,--output-files,--session,--record-idand--records-dirare gone; each value is derived, and no agent command in 1248 records passed any of them.Tried and withdrawn. For a day the branch also recorded any command with no registration (
dsagt-run -- <command>), under a rule that every computation leaves a record, and kept a copy of one-off scripts so they would replay. In 21 runs, 78 of the 81 records this produced were inspection one-liners that wrote no file, and they made up most of a reconstructed script's steps. How the reconstruction should cover a one-off shell step the pipeline depends on (an unpack, a move) stays open.3. The data-quality check runs on tables, once per file
Background. dsagt asks the agent to run AIDRIN, a data-quality tool for tables, before and after each step that changes a table, and to keep the reports.
Problem.
Change. Planned, one pull request (
1bf4eb2), after section 2.readiness_reports(path), returns the AIDRIN report on record for a file: the one made while the file had its present content, with the text the check printed, and the reports from before the file changed.4. One way for a skill's programs to become registered commands
Problem. A skill reaches a project four ways: dsagt installs three built-in skills when the project is created; the agent installs one from a catalog; the agent writes its own; or the agent registers a single program. Each way had its own code, and they did different things. The built-in skills' programs were registered from hand-written entries. A catalog skill's programs, and the agent's own, were registered only if the agent thought to do it. The skill's text kept telling the agent to run the programs directly (
python3 scripts/introspect.py), which leaves no record, and Claude did what the text said.Change. Planned, one pull request.
scan-directory, the one example command packaged inside dsagt, is removed with the code that supported packaged commands (8980171). It listed files, and taught the agent that reading the filesystem needs a registered command. The built-in skills' programs are the examples now.skills.register_skill_scripts, serves all four ways (c626ca7,53a12f3, part ofac257fa). It registers every runnable program under the skill'sscripts/, taking the parameters from a hand-written entry if there is one and otherwise from the program's own argument definitions. It rewrites the direct invocations in the skill's text to the recorded form, and replies to the agent with the lines to use. Registered commands and skills now share one folder,skills/: a registered command is a skill whose header names an executable. A test holds all four ways to the same result..claude/skills,.agents/skills), as a symbolic link per skill (c2c7fd4). As a real copy, a skill existed twice and the two drifted apart the moment either changed: under Claude the platform's copy went stale, and Codex kept them in step by editing both by hand, 14 times in one run. With a link there is one copy, so an edit to a skill, by the agent or by a person, is what every reader sees.save_code_specunder a name a script already has replies with that name instead of registering the script twice, and under an installed skill's name it is refused, since a code and a skill share one directory. The reply says to save the same spec under the existing name, which is how an argparse-derived spec gains its file roles and dependencies.search_skillsdescribes the catalogs it searches. It said it spanned installed skills, which every agent discovers natively and dsagt does not index.Overlap with the dataset-builder work (#43). That work places packaged commands in
src/dsagt/codes/, which this section removes, and #69 adds its checks as a skill whose program imports from the dsagt package. The rule now inCLAUDE.mdis: tools on the dsagt server own dsagt's own state; a packaged program that works on the user's files runs throughdsagt-runand imports nothing from dsagt. Under that rule the staleness check is a server tool and the rest is a skill registered through this section's function. Branch46-check-dataset-staleness-toolshows that shape as a proposal. The order of #52, #53, #69 against this section is the one open coordination question.5. Withdrawn: a check at the point where the agent issues a command
A Claude Code hook that put a bare python call under
dsagt-runwas built and validated, then removed. With section 4 in place the pipeline's steps are recorded without it; what it caught was inspection one-liners, which the conversation traces already hold. It covered one agent platform of five, edited the user's.claude/settings.json, and changed a command without the agent knowing. Codex, with no hook, recorded its pipeline steps from the instructions alone.6. The agent can find past runs; three duplicate tools go
Background. dsagt keeps several searchable collections for the agent: documents the user adds, registered commands, records of past runs, and memories.
Problem. The tool that lists collections left out dsagt's own, so the agent had no way to learn that past runs were searchable, and the search tool could not narrow a search to one command. Separately, the dsagt server offered
run_command,read_fileandhttp_request, which every agent platform already has. A command run through the server'srun_commandruns in the server's environment, not the user's, and leaves no record.Change. Planned, one pull request (
9855d1c).kb_list_collectionslists every collection with what it is for, the fields a search can filter on, and its size.kb_searchtakes a filter. The three tools are removed.install_dependenciesgoes with them, and so does the installsave_code_specran for a spec with dependencies: both installed into the dsagt server's own interpreter, which is not the environment the agent's shell runs a command in, and a spec's dependencies already reach the run through theuv run --withprefix in its stored line. The tool was called in none of 325 agent transcripts. The server has 17 tools.7. The server sees the user's Python environment under Codex and Cline
Problem. Codex and Cline start the dsagt server with only the environment variables written in their config file, not the ones from the shell the user launched from. The server then ran with a Python that lacked the packages in the user's activated environment.
Change. Planned, one pull request (
63a6c88).dsagt initanddsagt startcopy a fixed list of variables from the launching shell into that config (PATH,VIRTUAL_ENV,CONDA_PREFIX,PYTHONPATH, the library paths), plus any names the user lists undermcp.env_passthrough. A name that looks like a credential (*_KEY,*_TOKEN,*_SECRET) is refused, and a test holds that.The embedding backend's own credential follows the same rule. It read
EMBEDDING_API_KEY, thenLLM_API_KEY, thenOPENAI_API_KEY, and its base URL fromEMBEDDING_BASE_URLorOPENAI_BASE_URL; the server also readembedding.api_keyfrom the project config. TheOPENAI_names are the agent's LLM-provider credentials, which dsagt never reads. One name each now, from the shell or~/.config/dsagt/env.8.
dsagt initworks without gitProblem.
dsagt initdownloads the built-in skills from GitHub and needed thegitprogram to do it. A user who installs dsagt withpipmay not have git.Change. Planned, one pull request (
579a8d2). The skills are downloaded as an archive over HTTPS, and the exact commit is still written down beside them. git is used only as a fallback, for a private repository that the download is refused for.9. A resumed conversation is logged once
Background. dsagt reads the agent's conversation file from disk and logs each turn (prompt, tool calls, tokens) to a local MLflow store, so
dsagt infocan show what a session did and cost.Problem. Each
claude --continueorcodex exec resumestarts a new dsagt session on the same conversation file, and dsagt logged every earlier turn again under the new session. One 10-prompt walkthrough showed 54 turns, anddsagt infocounted the tokens several times over.Change. Open: #77. dsagt remembers which turns it has logged per conversation file, not per session, and leaves a turn that is still in progress for the next pass. Open: #76, a one-line fix to an integration test, found when the full test suite was run before the validation.
10. Documents and comments brought in line with the code
Problem. The README, the documentation site, and the comments in the code had drifted from what the code does, and from the project's writing rules.
Change. Planned, last. 18 false statements corrected, one test that checked nothing removed, the writing rules applied across the source, tests, site and walkthroughs, and the changelog entries for everything above.
A second pass on 2026-09-21 covered what the branch had since made false: the module docstrings for
observability.py,registry.pyandprovenance.py; the agent card, which advertised an embedder, a reranker, a tool count and a tool argument the code does not have; the per-operation check rule, which still told the agent to write reports toaudit/directly above the paragraph saying the report is the run's record; and the use-case catalog, which the site generates. It also removed two hand-tests describing an agent dsagt does not support, an API key committed totests/test_site_config.yaml, and a failing integration test whoseintegrationmarker had kept it out of the green run.The pass lists, for a later one and not changed here: four pieces of code nothing calls, nine alive only in their tests, eleven places where an error is caught and dropped,
dsagt start --agentwhich is declared and never read, anddsagt rm --allwhich is undocumented and untested.Upstream
datacard-generatorskill. Filled in as its own comments say, its template fails its own validator (41 errors), and agents stopped at its step "address findings" with findings left. Open: datacard-generator: the template validates when filled as its comments say; step 10 runs until the validator reports OK genesis-skills#16.noisy/folder as a side effect; one stale line in its skill's reference.Order
Sections 1 (first part) and 9 are open now as #74, #76 and #77. The rest is five pull requests, stacked because they change many of the same files (
provenance.py,skills.py,registry_tools.py, the agent's instructions), each based on the one before:dsagt-runand the record: sections 2 and 3, and the wrapper's one option.Each description ties its problem and fix to its commits and points back here for the background. The branch
comprehensive-reviewholds the same tree as the top of the stack.How the walkthroughs are judged changed on 2026-09-19. Three runs of each walkthrough on one tree showed a post-condition count moving by one or two between identical runs, and a third of the misses came from the headless harness itself (a turn ending a backgrounded command, nobody to answer a skill's second question). Mechanical checks of dsagt now fail a run; agent-dependent outcomes are reported as rates over three or more runs; what a headless session cannot show is not scored. The before-and-after table above predates that and is kept as the record of what started the work.
Tabled
Under consideration, not planned: an interface for a person, whose controls are the same tools the agent holds.