Skip to content
mist-medicalPublic

About

MIST: A simple and scalable end-to-end framework for 3D medical imaging segmentation.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

57 stars

Watchers

4 watching

Forks

Medical Imaging Segmentation Toolkit

GitHub Actions Workflow Status Docs GitHub Repo stars Coverage

MIST is a simple, scalable, end-to-end framework for 3D medical image segmentation. It handles everything from raw NIfTI files to trained models and evaluated predictions, with sensible defaults that work well out of the box and a configuration file for when you need more control.

Tip

Try MIST in your browser — no setup required. The end-to-end Colab demo runs the whole pipeline — analyze → preprocess → train → evaluate → predict → postprocess → rank → visualize — on a heart MRI dataset using a free Colab T4 GPU.

Open In Colab

Installation

CPU (default — runs the full pipeline: train, predict, evaluate, everything):

pip install mist-medical

NVIDIA GPU (recommended — adds DALI-accelerated data loading):

pip install "mist-medical[dali]"

Note

Upgrading from an older MIST release? pip install "mist-medical[train]" still works too — it's kept as an identical alias for [dali], so nothing in an existing script or notebook breaks.

AMD ROCm GPU (install a ROCm-enabled PyTorch build first — PyPI's default torch wheel has no ROCm support — matching your driver's ROCm version, then MIST on top with no extra needed; see Accelerator support for details):

pip install torch --index-url https://download.pytorch.org/whl/rocm6.4
pip install mist-medical

Development:

git clone https://github.com/mist-medical/MIST.git
cd MIST
pip install -e .            # CPU
pip install -e ".[dali]"    # NVIDIA GPU acceleration (recommended on CUDA)

Docker (NVIDIA GPU required, driver ≥ 525.x):

docker pull mistmedical/mist:latest

Quick Start

1. Prepare a dataset JSON describing your data:

{
  "task": "brats2023",
  "modality": "mr",
  "train-data": "/full/path/to/raw/data/train",
  "test-data": "/full/path/to/raw/data/validation",
  "mask": ["seg.nii.gz"],
  "images": {
    "t1": ["t1n.nii.gz"],
    "t2": ["t2w.nii.gz"],
    "tc": ["t1c.nii.gz"],
    "fl": ["t2f.nii.gz"]
  },
  "labels": [0, 1, 2, 3],
  "final_classes": { "WT": [1, 2, 3], "TC": [1, 3], "ET": [3] }
}

2. Run the full pipeline (analyze → preprocess → train → evaluate):

mist_run_all --data dataset.json \
             --numpy /path/to/numpy \
             --results /path/to/results

3. Run inference on new data:

mist_predict --models-dir /path/to/results/models \
             --config /path/to/results/config.json \
             --paths-csv /path/to/test.csv \
             --output /path/to/predictions

Key Features

  • Automatic configuration — analysis step determines target spacing, patch size, normalization, and foreground cropping from your data
  • Five-fold cross-validation by default, with custom fold assignment support
  • Multi-GPU training via PyTorch DDP on NVIDIA CUDA or AMD ROCm; uses all visible GPUs automatically
  • GPU-accelerated data loading via NVIDIA DALI during training, with a generic CPU-based data loader used automatically on AMD ROCm or CPU-only hardware — no NVIDIA-specific dependencies required to train at all
  • Sliding window inference with configurable overlap and patch blending
  • Test-time augmentation and multi-model ensembling at inference
  • Postprocessing with learnable, per-class morphological strategies
  • CPU training and inference — the full pipeline, mist_train included, runs on any machine, including Macs and laptops without an NVIDIA GPU

Supported Architectures

Model Key
nnU-Net nnunet
nnU-Net Pocket nnunet-pocket
MedNeXt (small / base / medium / large) mednext-small, mednext-base, mednext-medium, mednext-large
FMG-Net fmgnet
W-Net wnet
Swin UNETR (small / base / large) swinunetr-small, swinunetr-base, swinunetr-large

Change the architecture in config.json or pass --model <key> at the command line.

Supported Loss Functions

Loss Key Notes
Dice dice
Dice + Cross-Entropy dice_ce
clDice cldice Composite
Boundary Loss bl Composite
Generalized Surface Loss gsl Composite
Hausdorff Distance One-Sided hdos Composite
Volumetric SDDL volumetric_sddl Composite
Vessel SDDL vessel_sddl Composite

Composite losses blend a region-based term with a boundary/distance term, weighted by a scheduled alpha. Schedules: constant, linear, cosine.

Pipeline Commands

Command Description
mist_run_all Run the full pipeline end-to-end
mist_analyze Analyze dataset and generate config.json
mist_preprocess Preprocess images into NumPy arrays
mist_train Train models
mist_finalize Aggregate out-of-fold predictions into results.csv and run test-set inference
mist_predict Run inference on new data
mist_ensemble Combine predictions from multiple models via STAPLE, majority vote, or probability averaging
mist_evaluate Evaluate predictions against ground truth
mist_postprocess Apply postprocessing strategies
mist_rank Rank multiple evaluation result CSVs BraTS-style
mist_average_weights Average model weights across folds
mist_convert_msd Convert Medical Segmentation Decathlon datasets
mist_convert_csv Convert CSV-formatted datasets

Documentation

Full documentation, including configuration reference and advanced topics, is at mist-medical.readthedocs.io.

What's New

  • October 2026 — MIST 2.3.0, the first stable release — MIST drops its release-candidate tag. 2.3.0 rolls up everything below since 2.0.2-rc: CPU and AMD ROCm training, mist_finalize for per-fold distributed training, probability-level and parallel ensembling, and the removal of the ANTs dependency. pip install --upgrade mist-medical to get it, or pip install --upgrade "mist-medical[dali]" for DALI-accelerated data loading on NVIDIA GPUs (highly recommended).
  • September 2026 — CPU and AMD ROCm training support — the full pipeline, including mist_train, now runs on CPU-only machines and AMD ROCm GPUs via a new generic, pure-PyTorch data loader, auto-selected whenever NVIDIA DALI isn't the right fit for the detected hardware (or isn't installed). Communication backend (nccl/RCCL on ROCm, gloo on CPU) and data loader selection are both detected automatically and persisted to config.json. The train install extra is now called dali, to reflect that it gates DALI's CUDA acceleration specifically, not training capability in general — pip install mist-medical trains everywhere on its own; pip install "mist-medical[dali]" is recommended on top of that for NVIDIA GPU training. [train] is kept working as an alias for [dali] for anyone upgrading from an older release. AMP (BF16) is also now correctly scoped to hardware with real matrix acceleration — NVIDIA Ampere+ and AMD CDNA/RDNA3+ — rather than trusting torch.cuda.is_bf16_supported(), which reports available on older AMD RDNA1/2 GPUs despite having no BF16 acceleration at all.
  • September 2026 — mist_finalize for distributed per-fold training — mist_train now supports the one-fold-per-node HPC pattern: when an invocation's --folds/training.folds doesn't cover every configured fold, it writes out-of-fold predictions for the fold(s) it trained and prints a reminder instead of guessing at final results from a partial run. Run mist_finalize --results <dir> once, after every per-fold job sharing that results directory has finished, to aggregate every fold's held-out predictions into results.csv and run held-out test-set inference. mist_finalize is safe to re-run any time — each run recomputes from whatever's currently on disk.
  • September 2026 — ANTs dependency removed — MIST's image I/O, preprocessing, postprocessing, evaluation, analysis, and inference now run entirely on SimpleITK; antspyx is no longer a dependency at all. A six-stage migration, each stage verified against real trained models (pocket nnU-Net, FMGNet, SwinUNETR-small) across a full 2-GPU H100 pipeline run before shipping.
  • July 2026 — Probability-level ensembling — mist_predict --output-probs writes each model's final softmax probability volume alongside its discrete prediction; mist_ensemble --input-type probabilities averages probability volumes from separately trained models before a single argmax, preserving confidence information that STAPLE/majority vote discard.
  • July 2026 — Parallel ensembling — mist_ensemble now accepts --num-workers-ensemble to combine predictions across patients in parallel worker processes.
  • July 2026 — Interactive Colab demo — run the full MIST pipeline end-to-end on a free Colab GPU with the heart segmentation notebook: analysis, training, evaluation, inference, postprocessing, and BraTS-style ranking, all in the browser with no local setup.
  • June 2026 — Official Docker image — mistmedical/mist:latest ships a CUDA 12.8 build. Requires NVIDIA driver ≥ 525.x.
  • June 2026 — 2.0.1 release candidate — BF16 automatic mixed precision replaces FP16 throughout training and inference, reducing memory use and eliminating gradient loss scaling. Sliding-window inference gains a tunable sw_batch_size parameter. mist_rank adds pairwise Wilcoxon significance testing via --significance-csv. Several targeted memory-reduction fixes land across the inference stack.
  • June 2026 — Multi-model ensembling — mist_ensemble combines discrete NIfTI predictions from two or more separately trained models into a single consensus segmentation via STAPLE (--ensemble-backend staple, default) or majority vote (--ensemble-backend majority_vote). Works for single-class and multi-class label maps.
  • May 2026 — 2.0.0 release candidate — BraTS-style multi-strategy ranking (mist_rank), a structured postprocessing transform registry with LLM-readable metadata (describe_transforms), and full pathlib + PEP 585/604 modernization across the codebase.
  • April 2026 — CPU inference support — mist_predict now runs on any machine, including Macs and laptops without an NVIDIA GPU. Install with pip install mist-medical (no GPU required).
  • March 2026 — Resume training — interrupted runs can be continued from the last checkpoint with --resume, with atomic checkpointing to prevent corruption.
  • March 2026 — GPU-aware automatic patch size — the analysis step now derives the patch size from available GPU memory, so the default configuration is hardware-appropriate without manual tuning.
  • March 2026 — Transfer learning — initialize encoders from pretrained weights with --pretrained-weights, and average model weights across folds with mist_average_weights.
  • March 2026 — Better training defaults — AdamW optimizer and gradient clipping are now the defaults, with the clipping threshold exposed via grad_clip_norm in config.json.
  • September 2025 — BraTS 2025 adult glioma challenge @ MICCAI 2025 — MIST takes 3rd place (repeat).
  • November 2024 — MedNeXt models — small, base, medium, and large variants added (mednext-small, mednext-base, mednext-medium, mednext-large).
  • October 2024 — BraTS 2024 adult glioma challenge @ MICCAI 2024 — MIST takes 3rd place.

Citation

If you use MIST in your work, please cite:

@article{celaya2024mist,
  title   = {MIST: A Simple and Scalable End-To-End 3D Medical Imaging Segmentation Framework},
  author  = {Celaya, Adrian and others},
  journal = {arXiv preprint arXiv:2407.21343},
  year    = {2024}
}

@article{celaya2022pocketnet,
  title   = {PocketNet: A Smaller Neural Network For Medical Image Analysis},
  author  = {Celaya, Adrian and others},
  journal = {IEEE Transactions on Medical Imaging},
  doi     = {10.1109/TMI.2022.3224873},
  year    = {2022}
}

About

MIST: A simple and scalable end-to-end framework for 3D medical imaging segmentation.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

57 stars

Watchers

4 watching

Forks

Releases

Packages

Used by

Contributors

Languages