You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Part of #43. Depends on #45 (contract), #44 (structured pipeline output), #46 (checks), #47 (scaffolder), #48 (reference model).
The workflow itself, as a built-in skill in src/dsagt/skills/dataset-builder/. This is where the bulk of the feature is: DSAgt is infrastructure, not an agent, so a dataset builder is a workflow (a skill), not new MCP tools.
Scope
SKILL.md with a copyable progress checklist covering: gather evidence, derive the consumer contract, characterize the producer side, reconcile and write the contract, choose map-style versus iterable, generate the dataset code, derive and confirm the package name with the user, scaffold the package, generate the reference model and training loop when no user model exists, validate, and report.
Reference documents, kept out of SKILL.md so they cost no tokens until loaded:
references/layout_decision.md: the map-style versus iterable decision matrix keyed on on-disk layout (one file per sample, one large array, chunked formats such as zarr, HDF5, parquet; indexable in O(1); fits in memory; sharded or streamed). Includes two footguns that cause silent corruption: an IterableDataset without worker sharding yields every sample once per worker, and HDF5 or netCDF file handles opened before a fork produce corrupt reads in worker processes.
references/split_policy.md: group-aware and time-aware splitting, the leakage taxonomy (patient, simulation case, tokamak shot, crystal structure, augmented copies of one image, overlapping time windows), and the split manifest format. Random row-level splits leak whenever samples share a group, so group-aware splitting is an explicit prompt, not a default. The manifest (which ids in which split, plus the seed) is written to disk rather than recomputed at runtime from a hash that changes when a file is added.
references/model_inspection.md: how to read a training loop or model definition for the consumer contract, and the verify-by-execution requirement.
Two operating modes
Standalone. No execution records. Ask for the data root and characterize it with the built-in scan-directory code.
Pipeline-aware. Derive the input surface from the terminal outputs of the structured reconstruct_pipeline output (Structured output format for reconstruct_pipeline #44). Use the code_use collection and search_registry to see which transformations already ran, so the generated __getitem__ does not re-normalize data a pipeline step already normalized, and genuinely per-epoch work (augmentation, random windowing) is placed correctly.
Acceptance criteria
Frontmatter valid; description carries the trigger phrasings (write a DataLoader, prepare data for training, make a PyTorch Dataset, make this data AI-ready).
Both operating modes described, with the standalone path exercising scan-directory and the pipeline-aware path exercising the structured reconstruct_pipeline output.
Part of #43. Depends on #45 (contract), #44 (structured pipeline output), #46 (checks), #47 (scaffolder), #48 (reference model).
The workflow itself, as a built-in skill in
src/dsagt/skills/dataset-builder/. This is where the bulk of the feature is: DSAgt is infrastructure, not an agent, so a dataset builder is a workflow (a skill), not new MCP tools.Scope
SKILL.mdwith a copyable progress checklist covering: gather evidence, derive the consumer contract, characterize the producer side, reconcile and write the contract, choose map-style versus iterable, generate the dataset code, derive and confirm the package name with the user, scaffold the package, generate the reference model and training loop when no user model exists, validate, and report.Reference documents, kept out of
SKILL.mdso they cost no tokens until loaded:references/layout_decision.md: the map-style versus iterable decision matrix keyed on on-disk layout (one file per sample, one large array, chunked formats such as zarr, HDF5, parquet; indexable in O(1); fits in memory; sharded or streamed). Includes two footguns that cause silent corruption: anIterableDatasetwithout worker sharding yields every sample once per worker, and HDF5 or netCDF file handles opened before a fork produce corrupt reads in worker processes.references/split_policy.md: group-aware and time-aware splitting, the leakage taxonomy (patient, simulation case, tokamak shot, crystal structure, augmented copies of one image, overlapping time windows), and the split manifest format. Random row-level splits leak whenever samples share a group, so group-aware splitting is an explicit prompt, not a default. The manifest (which ids in which split, plus the seed) is written to disk rather than recomputed at runtime from a hash that changes when a file is added.references/model_inspection.md: how to read a training loop or model definition for the consumer contract, and the verify-by-execution requirement.Two operating modes
scan-directorycode.reconstruct_pipelineoutput (Structured output format forreconstruct_pipeline#44). Use thecode_usecollection andsearch_registryto see which transformations already ran, so the generated__getitem__does not re-normalize data a pipeline step already normalized, and genuinely per-epoch work (augmentation, random windowing) is placed correctly.Acceptance criteria
scan-directoryand the pipeline-aware path exercising the structuredreconstruct_pipelineoutput.