Skip to content

Epic: training-ready dataset builder (PyTorch) #43

Description

@ajtritt

Tracking issue for the training-ready dataset builder. No code lands here; the work is in the linked sub-issues.

Goal

Add a capability to DSAgt that turns a data directory (optionally one produced by a DSAgt curation pipeline) into a distributable Python package containing a PyTorch Dataset, a collate_fn, a split policy, and a build_dataloaders() factory, validated by executed checks rather than by inspection.

Two operating modes:

  1. Standalone. No DSAgt pipeline exists. The data root is supplied by the user and characterized with registered codes.
  2. Pipeline-aware. A DSAgt curation pipeline has run in this project. The builder derives its input surface from the recorded executions and knows which transformations have already been applied.

Framing

The feature is named for the dataset package, not for the DataLoader. A DataLoader is roughly five lines of configuration (batch_size, num_workers, pin_memory, sampler, drop_last). The difficulty is in the Dataset or IterableDataset, the collate_fn, and the split policy. Naming the feature after the loader points implementation effort at the easy half. Loader knobs become one configuration section of the generated output.

Core design: the sample contract

The pivot artifact is a sample contract: a persisted YAML file at <project>/dataset_contract.yaml declaring one sample's keys, dtypes, shapes, value ranges, semantics, the collation rule, and the split policy. Everything downstream derives from it, and it is what the user signs off on, so review happens against a readable spec rather than against generated code.

Two contracts get reconciled, and the Dataset is the adapter between them:

  • Consumer contract (model side): what forward() accepts.
  • Producer contract (data side): what is actually on disk.

The reconciliation list is the real design work: dtype casts, channel layout (NCHW versus NHWC), ownership of normalization (already applied by the pipeline, or applied in __getitem__), padding and ragged handling, label encoding.

Evidence hierarchy for the consumer contract

The builder uses the highest available source and falls through:

  1. An existing training loop or model definition in the user's repository.
  2. An existing Dataset the user wants replaced or improved.
  3. A reference paper, a model card, or a Hugging Face model id.
  4. Elicitation: ask the user to describe a canonical sample. At this tier there is no model to verify against, so the builder generates a reference model and training loop as part of the deliverable.

Static inspection is a hypothesis, not a conclusion. Keyword arguments, configuration-driven construction, dynamic shapes, and shape logic buried in a trainer all defeat static reading. Every statically inferred contract is verified by execution.

Validation as the deliverable

Generated code is accepted only after executed checks emitting JSON to audit/:

Check Catches
Contract ds[0] does not match declared keys, dtypes, shapes, ranges
Determinism Same seed produces a different sample
Worker equivalence num_workers=2 yields a different multiset than num_workers=0 (iterable duplication)
Split leakage Group ids appearing in more than one split; declared sizes wrong
Throughput Pathologically slow __getitem__
Model forward Batch fails to pass through model.forward()
Contract staleness The upstream pipeline changed after the contract was written

This is what separates the feature from a prompt that writes a Dataset.

Where the code goes

DSAgt is infrastructure, not an agent; capabilities are MCP services and workflows are skills. A dataset builder is a workflow, so the bulk is a built-in skill in src/dsagt/skills/, with its executable checks and scaffolding as built-in codes in src/dsagt/codes/.

The 20-tool MCP surface absorbs exactly one change, and it is a genuine infrastructure gap rather than workflow: reconstruct_pipeline renders only bash or Snakemake text, so the dependency graph it computes internally is not machine-readable by the caller.

Prior art

use_cases/fusion-fm/skills/xgc-ai-training/ is a hand-written, domain-specific instance of this exact artifact: check codes, a preprocessing step to npz, an XGCGraphDataset with a cached topology tensor and lazy per-item reads, a build_datasets() split factory, and a smoke test. It establishes the target artifact shape and serves as the end-to-end validation case.

Sub-issues

Issue Title Layer Depends on
#45 Sample-contract schema and persisted artifact design, spec none
#44 Structured output format for reconstruct_pipeline provenance.py, mcp/registry_tools.py none
#46 Built-in check-dataset code src/dsagt/codes/ #45, #44
#49 Built-in dataset-builder skill src/dsagt/skills/ #45, #44, #46, #47, #48
#47 Built-in scaffold-dataset-package code src/dsagt/codes/ #45
#50 Model-source carve-out in agent instructions dsagt_instructions.md none
#51 End-to-end validation against the XGC dataset shape use_cases/, docs/ #49, #47
#48 Reference model and training loop for the elicitation path src/dsagt/codes/ #45, #47

#44 is the only one that changes existing shipped behavior. #45 blocks #46, #47, #48, and #49, since all four consume the contract. #44, #45, #50 have no dependencies and can start immediately.

Out of scope for v1

  • Training loops.
  • Distributed samplers and multi-node orchestration.
  • An abstraction layer over PyTorch Lightning or Hugging Face Accelerate. A thin adapter is possible later.
  • TensorFlow and JAX.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions