Tracking issue for the training-ready dataset builder. No code lands here; the work is in the linked sub-issues.
Goal
Add a capability to DSAgt that turns a data directory (optionally one produced by a DSAgt curation pipeline) into a distributable Python package containing a PyTorch Dataset, a collate_fn, a split policy, and a build_dataloaders() factory, validated by executed checks rather than by inspection.
Two operating modes:
- Standalone. No DSAgt pipeline exists. The data root is supplied by the user and characterized with registered codes.
- Pipeline-aware. A DSAgt curation pipeline has run in this project. The builder derives its input surface from the recorded executions and knows which transformations have already been applied.
Framing
The feature is named for the dataset package, not for the DataLoader. A DataLoader is roughly five lines of configuration (batch_size, num_workers, pin_memory, sampler, drop_last). The difficulty is in the Dataset or IterableDataset, the collate_fn, and the split policy. Naming the feature after the loader points implementation effort at the easy half. Loader knobs become one configuration section of the generated output.
Core design: the sample contract
The pivot artifact is a sample contract: a persisted YAML file at <project>/dataset_contract.yaml declaring one sample's keys, dtypes, shapes, value ranges, semantics, the collation rule, and the split policy. Everything downstream derives from it, and it is what the user signs off on, so review happens against a readable spec rather than against generated code.
Two contracts get reconciled, and the Dataset is the adapter between them:
- Consumer contract (model side): what
forward() accepts.
- Producer contract (data side): what is actually on disk.
The reconciliation list is the real design work: dtype casts, channel layout (NCHW versus NHWC), ownership of normalization (already applied by the pipeline, or applied in __getitem__), padding and ragged handling, label encoding.
Evidence hierarchy for the consumer contract
The builder uses the highest available source and falls through:
- An existing training loop or model definition in the user's repository.
- An existing
Dataset the user wants replaced or improved.
- A reference paper, a model card, or a Hugging Face model id.
- Elicitation: ask the user to describe a canonical sample. At this tier there is no model to verify against, so the builder generates a reference model and training loop as part of the deliverable.
Static inspection is a hypothesis, not a conclusion. Keyword arguments, configuration-driven construction, dynamic shapes, and shape logic buried in a trainer all defeat static reading. Every statically inferred contract is verified by execution.
Validation as the deliverable
Generated code is accepted only after executed checks emitting JSON to audit/:
| Check |
Catches |
| Contract |
ds[0] does not match declared keys, dtypes, shapes, ranges |
| Determinism |
Same seed produces a different sample |
| Worker equivalence |
num_workers=2 yields a different multiset than num_workers=0 (iterable duplication) |
| Split leakage |
Group ids appearing in more than one split; declared sizes wrong |
| Throughput |
Pathologically slow __getitem__ |
| Model forward |
Batch fails to pass through model.forward() |
| Contract staleness |
The upstream pipeline changed after the contract was written |
This is what separates the feature from a prompt that writes a Dataset.
Where the code goes
DSAgt is infrastructure, not an agent; capabilities are MCP services and workflows are skills. A dataset builder is a workflow, so the bulk is a built-in skill in src/dsagt/skills/, with its executable checks and scaffolding as built-in codes in src/dsagt/codes/.
The 20-tool MCP surface absorbs exactly one change, and it is a genuine infrastructure gap rather than workflow: reconstruct_pipeline renders only bash or Snakemake text, so the dependency graph it computes internally is not machine-readable by the caller.
Prior art
use_cases/fusion-fm/skills/xgc-ai-training/ is a hand-written, domain-specific instance of this exact artifact: check codes, a preprocessing step to npz, an XGCGraphDataset with a cached topology tensor and lazy per-item reads, a build_datasets() split factory, and a smoke test. It establishes the target artifact shape and serves as the end-to-end validation case.
Sub-issues
| Issue |
Title |
Layer |
Depends on |
| #45 |
Sample-contract schema and persisted artifact |
design, spec |
none |
| #44 |
Structured output format for reconstruct_pipeline |
provenance.py, mcp/registry_tools.py |
none |
| #46 |
Built-in check-dataset code |
src/dsagt/codes/ |
#45, #44 |
| #49 |
Built-in dataset-builder skill |
src/dsagt/skills/ |
#45, #44, #46, #47, #48 |
| #47 |
Built-in scaffold-dataset-package code |
src/dsagt/codes/ |
#45 |
| #50 |
Model-source carve-out in agent instructions |
dsagt_instructions.md |
none |
| #51 |
End-to-end validation against the XGC dataset shape |
use_cases/, docs/ |
#49, #47 |
| #48 |
Reference model and training loop for the elicitation path |
src/dsagt/codes/ |
#45, #47 |
#44 is the only one that changes existing shipped behavior. #45 blocks #46, #47, #48, and #49, since all four consume the contract. #44, #45, #50 have no dependencies and can start immediately.
Out of scope for v1
- Training loops.
- Distributed samplers and multi-node orchestration.
- An abstraction layer over PyTorch Lightning or Hugging Face Accelerate. A thin adapter is possible later.
- TensorFlow and JAX.
Tracking issue for the training-ready dataset builder. No code lands here; the work is in the linked sub-issues.
Goal
Add a capability to DSAgt that turns a data directory (optionally one produced by a DSAgt curation pipeline) into a distributable Python package containing a PyTorch
Dataset, acollate_fn, a split policy, and abuild_dataloaders()factory, validated by executed checks rather than by inspection.Two operating modes:
Framing
The feature is named for the dataset package, not for the
DataLoader. ADataLoaderis roughly five lines of configuration (batch_size,num_workers,pin_memory,sampler,drop_last). The difficulty is in theDatasetorIterableDataset, thecollate_fn, and the split policy. Naming the feature after the loader points implementation effort at the easy half. Loader knobs become one configuration section of the generated output.Core design: the sample contract
The pivot artifact is a sample contract: a persisted YAML file at
<project>/dataset_contract.yamldeclaring one sample's keys, dtypes, shapes, value ranges, semantics, the collation rule, and the split policy. Everything downstream derives from it, and it is what the user signs off on, so review happens against a readable spec rather than against generated code.Two contracts get reconciled, and the
Datasetis the adapter between them:forward()accepts.The reconciliation list is the real design work: dtype casts, channel layout (NCHW versus NHWC), ownership of normalization (already applied by the pipeline, or applied in
__getitem__), padding and ragged handling, label encoding.Evidence hierarchy for the consumer contract
The builder uses the highest available source and falls through:
Datasetthe user wants replaced or improved.Static inspection is a hypothesis, not a conclusion. Keyword arguments, configuration-driven construction, dynamic shapes, and shape logic buried in a trainer all defeat static reading. Every statically inferred contract is verified by execution.
Validation as the deliverable
Generated code is accepted only after executed checks emitting JSON to
audit/:ds[0]does not match declared keys, dtypes, shapes, rangesnum_workers=2yields a different multiset thannum_workers=0(iterable duplication)__getitem__model.forward()This is what separates the feature from a prompt that writes a
Dataset.Where the code goes
DSAgt is infrastructure, not an agent; capabilities are MCP services and workflows are skills. A dataset builder is a workflow, so the bulk is a built-in skill in
src/dsagt/skills/, with its executable checks and scaffolding as built-in codes insrc/dsagt/codes/.The 20-tool MCP surface absorbs exactly one change, and it is a genuine infrastructure gap rather than workflow:
reconstruct_pipelinerenders only bash or Snakemake text, so the dependency graph it computes internally is not machine-readable by the caller.Prior art
use_cases/fusion-fm/skills/xgc-ai-training/is a hand-written, domain-specific instance of this exact artifact: check codes, a preprocessing step to npz, anXGCGraphDatasetwith a cached topology tensor and lazy per-item reads, abuild_datasets()split factory, and a smoke test. It establishes the target artifact shape and serves as the end-to-end validation case.Sub-issues
reconstruct_pipelineprovenance.py,mcp/registry_tools.pycheck-datasetcodesrc/dsagt/codes/dataset-builderskillsrc/dsagt/skills/scaffold-dataset-packagecodesrc/dsagt/codes/dsagt_instructions.mduse_cases/,docs/src/dsagt/codes/reconstruct_pipeline#44 Structured output format forreconstruct_pipelinecheck-datasetcode #46 Built-incheck-datasetcodescaffold-dataset-packagecode #47 Built-inscaffold-dataset-packagecodedataset-builderskill #49 Built-indataset-builderskill#44 is the only one that changes existing shipped behavior. #45 blocks #46, #47, #48, and #49, since all four consume the contract. #44, #45, #50 have no dependencies and can start immediately.
Out of scope for v1