Skip to content

How do pyramids change in V_eta? #1007

Description

@stevevanhooser

NDI currently defines three unrelated families of "pyramid" documents. They were each built to satisfy a specific viewer / ingest path and don't share a common schema, common member-file convention, or common depends_on topology. As we plan V_eta, we should decide whether to reshape them into a single family (with a shared traversal and a shared viewer contract), or to formalize a common subset of fields the viewers can all rely on while leaving the specifics per family.

This issue collects what each of the three families holds today so the reshaping conversation has one place to work from.


1. Spatial transcriptomics (spatialGeneExpressionPyramid and friends)

Three coordinating documents, one pyramid parent + two per-bin-size children.

spatialGeneExpressionPyramid — the parent. One per (subject, geneList).
class_version: 1. depends_on: subject_id, geneList_id. superclasses: base, data/geneExpression.
Key fields: label, chip_serial, pipeline_version, bin_sizes (e.g. [1,2,4,8,16,32]), base_pixel_size_x/y, pixel_size_units, origin_x/y, extent_x/y, tile_rows, tile_columns, index_order (row-major), origin_corner (upper-left), byte_order.
Files: gene_totals.tsv (one summary file on the parent).

spatialGeneExpressionTiles — one per bin size. Attached to the parent by spatialGeneExpressionPyramid_id.
depends_on: spatialGeneExpressionPyramid_id, subject_id, source_file_id.
Key fields: bin_size, pixel_size_x/y, dimension_order (YXG), dimension_labels (height,width,gene), dimension_size, dimension_scale, tile_size_x/y_bins, n_tiles_stored, data_type_gene_index/count/offset/coordinate, tile_compression (none), tile_format_version, tile_index_origin (1).
Files: legacy NAME# series — "tile.bin_#" in file_list.
Sparse-by-design: a gene pyramid does not write empty tiles, so the "walk NAME_1..N until a hole" scheme it used to rely on gave silent truncation. See NDI-matlab#956 for the ingest fix.

spatialGeneExpressionCells — one per bin size. Cell segmentation of the same tissue plane.
depends_on: same three as Tiles.
Key fields: n_cells, segmentation_method, segmentation_dilation, coordinate_units, contours_present, contour_reference (centroid), n_vertices_per_cell, data_type_vertex/offset, contour_format_version.
Files: cells.tsv, contours.bin (regular files, not a series).

Overall pattern: level = bin_size, dimensions are 2D (Y, X) plus a gene axis, no explicit reduction ladder (each level is a sum over its own tile size, not a downsample of a lower level).

2. Lightsheet OME-Zarr (lightsheetZarrPyramid + lightsheetZarrLevel)

Two documents, one parent + one per level. Built on this branch for the SmartSPIM ingest / napari viewer path.

lightsheetZarrPyramid — the parent. One per Zarr group.
class_version: 1. depends_on: subject_id, element_id, source_file_id. superclasses: base.
Key fields: label, pyramid_name, pyramid_type, axes_order (tczyx), axes_units, channel_names, n_channels, shape_level0, voxel_size_level0, voxel_size_units, translation_level0, dtype, pipeline_version, byte_order.
Files: none on the parent — the level docs carry the bytes.

lightsheetZarrLevel — one per level in the ladder. Attached by lightsheetZarrPyramid_id.
depends_on: lightsheetZarrPyramid_id, subject_id, source_file_id.
Key fields: level, reduction_function (none for raw level 0; mean or max for reduced levels), axes_order, shape, chunks, chunk_grid, n_chunks_stored, chunk_index_origin (1), chunk_order (C), dtype, fill_value, byte_order, codec (raw, blosc, ...), codec_params, voxel_size, voxel_size_units, translation.
Files: DID new-style file series — chunk.bin declared under files.file_series, with members chunk.bin_1..N produced from the pyramid ingest. Sparse-safe by construction: the manifest names each present slot's uid, so absent slots are structural, not a "first hole = end" walk.

Level ladder holds two reduction functions in parallel (mean + max), sharing a single raw level-0 (reduction_function=none). The viewer picks one reduction via --reduction; default now is mean.

3. Electrophysiology (pyraview)

One document per element, one file series conceptually but stored as a fixed set of level files.

pyraview — one per element.
class_version: 1. depends_on: element_id. superclasses: epochclocktimes, data/filter.
Key fields: label, nativeRate, nativeStartTime, channels, dataType, decimationLevels, decimationSamplingRates, decimationStartTimes.
Files: fixed list — level1.bin, level2.bin, ..., level10.bin in file_list. Not a series (no #).

Overall pattern: level = decimation step, dimensions are 1D (time) × channels, ladder length pinned to 10, each level stored as one contiguous binary blob (no chunking). Extending or shrinking the ladder is a schema change, not a data change.


What actually differs across the three

Aspect Gene expression Lightsheet Zarr Pyraview
Parent + children? Yes (1 + 2 per bin) Yes (1 + N levels) No, single doc
Level = ? bin_size (aggregation) level (2× downsample) decimationLevels[i]
Reduction Sum over tile bins mean / max / none Filtered decimation
Multiple reductions? No Yes (mean and max share raw) No
Level ladder length Fixed in schema (bin_sizes) Data-driven per pyramid Fixed at 10 in schema
Chunk / tile mechanism tile.bin_# NAME# series chunk.bin file_series levelN.bin fixed files
Sparse levels supported Yes (must, gene tiles are sparse) Yes (structural, via manifest) No (level file is dense)
Chunk grid / index origin recorded No Yes (chunk_grid, chunk_index_origin) N/A
Codec metadata tile_compression (none only) codec + codec_params (raw, blosc, ...) None
Voxel/pixel metadata Per bin: pixel_size_x/y Per level: voxel_size + voxel_size_units nativeRate on parent only
dtype Split across data_type_* fields Single dtype field dataType field
Byte order Yes (byte_order) Yes Not recorded
depends_on topology Parent ← Tiles ← subject, geneList, source_file Parent ← Level ← subject, element, source_file element only
Viewer GEF / spatial-genomics napari napariViewLightsheet pyraview GUI

Questions V_eta should answer

  1. One family or three? Do the three collapse into a shared pyramid + pyramidLevel pair (with a per-family sub-property block), or do we keep three families and standardize only the field NAMES that viewers rely on?
  2. What's the shared subset a viewer can always ask for? Candidates: label, axes_order, shape per level, voxel_size per level, voxel_size_units, dtype, byte_order, reduction_function, chunk_grid / member layout, pipeline_version. Everything else is per family.
  3. File-member convention. New-style file_series (with a manifest) is what lightsheet uses today; gene expression is still on NAME#. Pyraview isn't a series at all. Should V_eta migrate everyone to file_series and retire NAME#?
  4. Reduction ladder. Lightsheet allows multiple reductions sharing a raw level 0. Gene expression's "reduction" is a bin-size aggregation with no shared root. Pyraview's is a filter chain. A shared field reduction_function with a per-family enum would let viewers switch reductions uniformly; do we want to require it?
  5. Level identity. Gene expression stores levels as separate documents indexed by bin_size; lightsheet indexes by level; pyraview indexes by array position in a fixed schema. If levels become documents everywhere, pyraview loses its "10 fixed files" simplicity. Is that a good trade?
  6. Class version. All three are class_version: 1 today. Any schema rework is a class_version: 2 bump on the affected documents. Should V_eta ship an upgrade path (definition + a converter) rather than a hard bump?

Cross-refs: NDI-matlab#956 (gene pyramid sparse-tile walk bug the new-style series retires), the current lightsheet ingest / viewer PR (#979), and DID-matlab#183/#188 (series-member conventions the lightsheet family is built on).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions