C++20 • CUDA + cuDNN • YOLO models • TensorFlow event metrics • Apache 2.0
A compact C++20 neural-network engine for CUDA-accelerated object detection.
Train, validate, inspect, and run YOLO-style models without hiding the machinery behind a framework.
Why PixieNN? · Demo · Models · Quick start · Training · Architecture
PixieNN is a modern rethinking of the ideas behind Darknet: small enough to understand, direct enough to debug, and capable of running the complete object-detection workflow in native code.
| Native performance | C++20, CUDA, cuDNN, and OpenBLAS execution paths—without a Python runtime in the training or inference loop. |
| The whole workflow | Training, validation, checkpointing, TensorFlow .tfevents data, image inference, non-max suppression, and GeoJSON predictions. |
| Readable experiments | Human-editable YAML describes model graphs, optimizer settings, augmentation, datasets, and learning-rate schedules. |
| Reproducible runs | GPU/data preflight checks, isolated run directories, metadata, logs, safe restart behavior, and explicit checkpoint resume. |
PixieNN is under active development. Its goal is not to impersonate a mature Python ecosystem; it is to offer a focused, inspectable native engine where model behavior can be traced all the way down to the kernels.
YOLOv3-tiny inference: bounding boxes are rendered to JPEG while the same detections are exported as GeoJSON.
The repository includes model graphs and runnable presets spanning small experiments through larger detectors.
| Family | Included preset | Intended use |
|---|---|---|
| CenterNet | centernet-smoke-voc, centernet-tiny-voc, centernet-kitti, centernet-coco |
Anchor-free VOC, KITTI, and COCO training |
| YOLOv1 | yolov1-tiny |
Compact VOC training and architecture experiments |
| YOLOv2 | yolov2 |
YOLOv2-style anchor-based VOC smoke tests |
| YOLO Nano | yolo-nano |
Minimal detector experiments on VOC |
| YOLOv3 Tiny | yolov3-tiny, yolov3-tiny-voc |
Fast COCO smoke tests or full VOC training |
| YOLOv3 | yolov3 |
Larger COCO model graph |
| YOLOv7-style | yolov7 |
Larger COCO training preset |
| ResNet-18 | model definition included | Classification and layer coverage |
Note
Model definitions describe PixieNN graphs and training presets. They should not be read as claims of published reference-paper accuracy. Reproducible benchmark checkpoints and PR curves are a project milestone, not a fabricated checkbox.
The CenterNet head provides a deliberately different detector for comparison with the YOLO families. It treats objects as center points and predicts a per-class center heatmap, normalized box width and height, and a fractional center offset. This removes anchor configuration from the experiment and makes the learned heatmaps directly inspectable.
Use centernet-smoke-voc to verify the complete target, loss, checkpoint, and
decode pipeline on a small network. Use centernet-tiny-voc for a real VOC
experiment. The current CUDA implementation performs the CenterNet head math on
the host while the convolutional backbone remains CUDA-accelerated; it favors a
clear, testable reference implementation over peak head throughput.
For a smaller real-world alternative to COCO, download and prepare KITTI's object-detection training set in Darknet layout:
./shell/kitti.sh
./shell/train-model.sh centernet-kitti --fresh --verify-dataRead CenterNet: Objects as Glowing Points for a visual, beginner-friendly tour of heatmaps, center offsets, box reconstruction, and how the anchor-free approach differs from YOLO.
git clone https://github.com/trieck/pixienn.git
cd pixienn
cmake -S . -B build \
-DUSE_CUDA=ON \
-DCMAKE_BUILD_TYPE=Release
cmake --build build --parallelThe main executables are written to build/bin/:
pixienn— image inferencepixienn-train— model training and validationpixienn-test— the native test suite
Run the tests with:
./build/bin/pixienn-test --gtest_brief=1Download the original YOLOv3-tiny Darknet weights:
mkdir -p resources/weights
curl -L https://pjreddie.com/media/files/yolov3-tiny.weights \
-o resources/weights/yolov3-tiny.weightsThen run PixieNN from the repository root:
./build/bin/pixienn \
--confidence=0.20 \
--nms=0.40 \
resources/cfg/yolov3-tiny-cfg.yml \
resources/images/dog.jpgThis writes predictions.jpg and predictions.geojson to the current directory.
GeoJSON is an intentional visualization format in PixieNN, not just a generic JSON export. It lets the detection boxes be loaded as a vector layer over the original, non-georeferenced image in QGIS. Each prediction is a polygon whose attributes include the class, confidence, and batch ID, so boxes can be inspected, filtered, styled, and compared without flattening them into the rendered JPEG.
- Add the inference image to QGIS as a raster layer.
- Add
predictions.geojsonas a vector layer. - Keep both layers in their local, non-georeferenced image coordinate space.
- Style or filter the vector layer using its
classandconfidenceattributes.
Image coordinates normally begin at the upper-left and increase downward. PixieNN writes negative GeoJSON Y coordinates so the polygons align with the way QGIS displays an unreferenced raster in its Cartesian canvas. The export is therefore intended for local image inspection; its coordinates are pixel-space geometry, not longitude and latitude.
The training wrappers standardize the parts of long-running GPU jobs that are easy to get wrong: executable selection, CUDA linkage, input manifests, stale output, logs, metadata, checkpoints, locking, and resume behavior.
List every available preset:
./shell/train-model.sh --listStart a clean run after verifying every training and validation sample:
./shell/train-model.sh yolov7 --fresh --verify-dataResume the latest complete checkpoint:
./shell/train-model.sh yolov7 --resumeRun all configured models sequentially:
./shell/train-all-models.sh --fresh --verify-dataThe wrappers deliberately reject CPU-only binaries. A run is organized for both humans and tooling:
runs/
├── archive/ # complete older runs preserved by --fresh
│ └── yolov7-<timestamp>/
└── yolov7/
├── backup/ # rolling checkpoints
├── run-metadata.txt # command, host, GPU, revision, and start time
├── training.log # captured console output
├── tensorboard.log # TensorBoard server output
├── tensorboard.pid # automatically started server PID
├── events.out.tfevents.*
└── *.weights # primary/final weights
The wrapper starts TensorBoard automatically and prints a clickable URL:
TensorBoard: http://localhost:6006/
With --fresh, it stops an existing TensorBoard instance on that port and
removes this model's previous event files before starting the clean run. Set
PIXIENN_TENSORBOARD_PORT to use a different port.
PixieNN currently reports training/validation loss, IoU, recall, micro-averaged F1, and mAP at IoU 0.50. Do not confuse the latter with COCO's stricter mAP averaged from IoU 0.50 through 0.95.
For preset overrides, data-manifest rules, dry runs, locking, and cleanup semantics, read the training guide.
The repository includes a local React dashboard for TensorFlow .tfevents
protocol-buffer data,
run metadata, and checkpoints. Start it from the repository root:
cd monitor
npm install
npm run devOpen http://localhost:4173. The dashboard refreshes
every 2.5 seconds and shows the current run status, optimizer step, loss,
learning rate, latest checkpoint, metadata, full-run loss trace, and
auto-scaled scalar cards sourced from the event files. The charts span the
entire run automatically; the renderer keeps its display resolution bounded as
the event file grows. The average-loss chart also supports exact recent-step
windows of 10,000, 2,000, or 500 optimizer steps. Use the run selector to inspect
another directory under runs/.
The monitor reads local event/metadata files through its local-only API; it
does not upload logs or expose arbitrary filesystem paths. TensorBoard remains
available at the separate URL printed by the training wrapper, normally
http://localhost:6006/.
flowchart LR
A[Dataset manifests] --> C[Experiment config]
B[Model graph + hyperparameters] --> C
C --> D[pixienn-train]
D --> E[Checkpoints]
D --> F[TensorFlow .tfevents data]
E --> G[pixienn inference]
G --> H[Annotated JPEG]
G --> I[GeoJSON detections]
resources/cfg/binds a model to weights, labels, and train/validation manifests.resources/models/defines the layer graph and training hyperparameters.include/contains the engine, CUDA layers, optimizers, metrics, and data pipeline.src/provides the CLI entry points and concrete implementation units.shell/contains reproducible training and dataset utilities.tests/exercises layers, training behavior, metrics, serialization, and scripts.
| Core | GPU acceleration | Image and configuration |
|---|---|---|
| CMake 3.15+ | NVIDIA CUDA Toolkit | OpenCV 4.5.4+ |
| C++20 compiler | cuDNN 8+ | LibTIFF |
| Boost 1.74+ | Compatible NVIDIA driver | yaml-cpp |
| OpenBLAS | nlohmann/json 3.10.5+ | |
| Protobuf 3.12.4+ | GLib and HarfBuzz |
Cairo and Pango are optional visualization dependencies. CUDA is optional at build time, but it is required by the guarded training wrappers documented above.
A run is split into two YAML files so datasets and model internals remain independently reusable:
# resources/cfg/<experiment>-cfg.yml
configuration:
model: ../models/<model>.yml
weights: ../weights/<checkpoint>.weights
labels: ../data/<dataset>/labels.txt
training:
train-images: ../data/<dataset>/train.txt
train-labels: ../data/<dataset>/labels/train
val-images: ../data/<dataset>/val.txt
val-labels: ../data/<dataset>/labels/valThe model YAML controls layers, batch sizing, augmentation, validation cadence, optimizer parameters, learning-rate policy, checkpointing, and early stopping. Start with an included preset; then change one variable at a time and preserve the generated run metadata.
Mosaic augmentation is opt-in and combines four training images before the
normal detector target builder runs. Add it under augmentation in a model
YAML:
augmentation:
enabled: True
flip: True
jitter: 0.2
saturation: 1.5
exposure: 1.5
hue: 0.1
mosaic:
enabled: True
probability: 0.5The probability controls how often a training sample is built from four source images. Validation never uses Mosaic because validation loaders do not create an augmenter.
The next meaningful milestones are evidence, not feature-count theater:
- publish reproducible checkpoints and precision–recall curves;
- add standards-compliant COCO mAP50–95 evaluation;
- document GPU throughput and memory benchmarks;
- expand end-to-end training regression coverage;
- continue simplifying the path from YAML to kernel execution.
Bug reports, focused pull requests, reproducible training observations, and new tests are welcome. When reporting training behavior, include the model/config YAML, GPU model, command line, relevant run-metadata.txt, and a short event-file export whenever possible.
PixieNN was inspired by Darknet's directness and its enduring contribution to real-time object detection. Source files are distributed under the Apache License 2.0.
If an understandable native training engine is useful to you, give PixieNN a ⭐ and help test it on new hardware and datasets.
