Skip to content

Built-in dataset-builder skill #49

Description

@ajtritt

Part of #43. Depends on #45 (contract), #44 (structured pipeline output), #46 (checks), #47 (scaffolder), #48 (reference model).

The workflow itself, as a built-in skill in src/dsagt/skills/dataset-builder/. This is where the bulk of the feature is: DSAgt is infrastructure, not an agent, so a dataset builder is a workflow (a skill), not new MCP tools.

Scope

SKILL.md with a copyable progress checklist covering: gather evidence, derive the consumer contract, characterize the producer side, reconcile and write the contract, choose map-style versus iterable, generate the dataset code, derive and confirm the package name with the user, scaffold the package, generate the reference model and training loop when no user model exists, validate, and report.

Reference documents, kept out of SKILL.md so they cost no tokens until loaded:

  • references/layout_decision.md: the map-style versus iterable decision matrix keyed on on-disk layout (one file per sample, one large array, chunked formats such as zarr, HDF5, parquet; indexable in O(1); fits in memory; sharded or streamed). Includes two footguns that cause silent corruption: an IterableDataset without worker sharding yields every sample once per worker, and HDF5 or netCDF file handles opened before a fork produce corrupt reads in worker processes.
  • references/split_policy.md: group-aware and time-aware splitting, the leakage taxonomy (patient, simulation case, tokamak shot, crystal structure, augmented copies of one image, overlapping time windows), and the split manifest format. Random row-level splits leak whenever samples share a group, so group-aware splitting is an explicit prompt, not a default. The manifest (which ids in which split, plus the seed) is written to disk rather than recomputed at runtime from a hash that changes when a file is added.
  • references/model_inspection.md: how to read a training loop or model definition for the consumer contract, and the verify-by-execution requirement.

Two operating modes

  • Standalone. No execution records. Ask for the data root and characterize it with the built-in scan-directory code.
  • Pipeline-aware. Derive the input surface from the terminal outputs of the structured reconstruct_pipeline output (Structured output format for reconstruct_pipeline #44). Use the code_use collection and search_registry to see which transformations already ran, so the generated __getitem__ does not re-normalize data a pipeline step already normalized, and genuinely per-epoch work (augmentation, random windowing) is placed correctly.

Acceptance criteria

  • Frontmatter valid; description carries the trigger phrasings (write a DataLoader, prepare data for training, make a PyTorch Dataset, make this data AI-ready).
  • Both operating modes described, with the standalone path exercising scan-directory and the pipeline-aware path exercising the structured reconstruct_pipeline output.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions