Skip to content

Structured output format for reconstruct_pipeline #44

Description

@ajtritt

Part of #43.

provenance.reconstruct_pipeline renders bash or Snakemake text. load_pipeline_records and build_dependency_graph (src/dsagt/provenance.py) already produce the underlying structure, but nothing escapes the function, so a caller wanting file lists has to parse shell comments out of render_bash output.

The dataset builder needs the DAG structurally: terminal outputs (files no downstream step consumes) are the dataset's input surface, and that derivation should be mechanical rather than reconstructed from a rendered script.

This is the only change in the epic that touches existing shipped behavior.

Scope

  • Add a JSON format to reconstruct_pipeline returning the records, the dependency graph, and the derived terminal outputs.
  • Add the format to the MCP tool's format enum in src/dsagt/mcp/registry_tools.py (currently ["bash", "snakemake"]).
  • Add terminal-output derivation: output files that appear in no other record's input files.

Acceptance criteria

  • Unit tests covering terminal-output derivation, including a diamond-shaped DAG and a pipeline with several independent leaves.
  • Existing bash and Snakemake output unchanged.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions