Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
23 changes: 23 additions & 0 deletions .dockerignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,23 @@
# Keep the build context small and avoid setup.py's dirty-tree assert.
.git
.github
.venv
venv
**/__pycache__
**/*.pyc
**/*.pyo
**/.DS_Store
**/.idea
**/.vscode
**/.cache
**/cmake-build-*/
build/
dist/
*.egg-info/
deep_ep.egg-info/
*.so
coredump/
compile_commands.json
.clangd
.sisyphus/
AGENTS.md
12 changes: 12 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -80,6 +80,18 @@ bash install.sh

Then import `deep_ep` in your Python project.

### Docker

A minimal PyTorch CUDA devel image that builds and installs DeepEP is under [`docker/`](docker/):

```bash
git submodule update --init --recursive
docker build --platform=linux/amd64 -f docker/Dockerfile -t deepep:local .
docker run --rm -it --gpus all --ipc=host deepep:local
```

See [docker/README.md](docker/README.md) for base-image overrides and limits (RDMA, multi-node, and host NVIDIA requirements).

### Development and tests

```bash
Expand Down
51 changes: 51 additions & 0 deletions docker/Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
# Official DeepEP image: PyTorch + CUDA devel toolkit + NCCL + built deep_ep.
#
# Build from the repository root (DeepJIT submodule must already be checked out):
# git submodule update --init --recursive
# docker build -f docker/Dockerfile -t deepep:local .
#
# Run on a Linux host with NVIDIA Container Toolkit:
# docker run --rm -it --gpus all --ipc=host deepep:local

# CUDA Toolkit 13.1+ is required by DeepEP's DeepJIT path; use a matching devel image
# so nvcc remains available at runtime for JIT compilation.
ARG BASE_IMAGE=pytorch/pytorch:2.14.1-cuda13.2-cudnn9-devel
FROM ${BASE_IMAGE}

LABEL org.opencontainers.image.title="DeepEP" \
org.opencontainers.image.description="DeepEP built on the official PyTorch CUDA devel image" \
org.opencontainers.image.source="https://github.com/deepseek-ai/DeepEP"

ENV DEBIAN_FRONTEND=noninteractive \
PYTHONUNBUFFERED=1 \
# Keep the toolkit discoverable for DeepJIT at import/runtime.
CUDA_HOME=/usr/local/cuda \
PATH=/usr/local/cuda/bin:${PATH}

# Minimal host packages beyond the PyTorch devel image.
# build-essential provides a C++20-capable toolchain on current Ubuntu bases.
RUN apt-get update && apt-get install -y --no-install-recommends \
build-essential \
ca-certificates \
git \
ninja-build \
&& rm -rf /var/lib/apt/lists/*

WORKDIR /opt/DeepEP

# Expect a clean build context that already contains third-party/deep_jit sources.
COPY . .

RUN test -f third-party/deep_jit/include/deep_jit/python_api.hpp \
|| (echo "DeepJIT submodule missing. Run: git submodule update --init --recursive" >&2 && exit 1)

# Match the README floor. --no-deps avoids pulling an older torch pin transitively.
# find_pkgs / EP_NCCL_ROOT_DIR will resolve this pip package at build and import time.
RUN python -m pip install --no-cache-dir --upgrade pip setuptools wheel \
&& python -m pip install --no-cache-dir "nvidia-nccl-cu13>=2.32.3" --no-deps \
&& python -m pip install --no-cache-dir numpy \
&& bash install.sh \
&& python -c "import deep_ep; print('deep_ep', deep_ep.__version__)"

WORKDIR /workspace
CMD ["bash"]
64 changes: 64 additions & 0 deletions docker/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,64 @@
# DeepEP Docker image

Minimal image based on the official PyTorch CUDA **devel** image, with DeepEP and its NCCL floor installed.

## Requirements

- Docker with [NVIDIA Container Toolkit](https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html) on a Linux host
- NVIDIA Hopper (or newer) GPU for running DeepEP kernels
- Repository checkout with the DeepJIT submodule initialized

The image installs the host extension at build time. DeepJIT still needs `nvcc` from the CUDA toolkit at runtime, which is why the base image is `*-devel` rather than `*-runtime`.

## Build

From the repository root:

```bash
git submodule update --init --recursive
# On Apple Silicon / other non-amd64 hosts, pin the platform to match the NVIDIA image:
docker build --platform=linux/amd64 -f docker/Dockerfile -t deepep:local .
```

Optional base-image override (must remain a CUDA 13.1+ devel image with PyTorch ≥ 2.10):

```bash
docker build --platform=linux/amd64 -f docker/Dockerfile \
--build-arg BASE_IMAGE=pytorch/pytorch:2.14.1-cuda13.2-cudnn9-devel \
-t deepep:local .
```

## Run

```bash
docker run --rm -it --gpus all --ipc=host deepep:local
```

`--ipc=host` (or a larger `--shm-size`) is recommended for PyTorch multiprocessing and multi-process tests.

Smoke check inside the container:

```bash
python -c "import deep_ep, torch; print(deep_ep.__version__, torch.cuda.is_available())"
```

Example single-node test (needs a visible GPU):

```bash
python /opt/DeepEP/tests/ep/test_ep.py
```

## What this image includes

| Component | Source |
| --- | --- |
| PyTorch + CUDA toolkit (`nvcc`) | `pytorch/pytorch:*-cuda13.2-cudnn9-devel` |
| NCCL ≥ 2.32.3 | `pip install nvidia-nccl-cu13>=2.32.3 --no-deps` |
| DeepEP | built with `bash install.sh` into the image Python env |
| NumPy | test dependency |

## Limits

- Inter-node RDMA / multi-node tests need matching host IB/RoCE setup and usually `--network=host` plus device mounts for RDMA; those are site-specific and are not baked into this Dockerfile.
- Building the image on macOS/arm64 without an NVIDIA GPU can produce an amd64 image via buildx, but you still need a Linux NVIDIA host to run DeepEP.
- The Docker build context excludes `.git`, so the installed package version suffix is typically `+local` (see `setup.py`).