Turn raw survey imagery or video into a curated, labeled DuckDB database, then out again as training data. Described in arXiv:2609.25500 (see Citing this work). Five stages:
| Stage | Commands | What it does |
|---|---|---|
| Import | import yolo, import voc |
an existing labeled dataset into a database |
| Generate | infer images, infer video |
pixels + detections into a database |
| Enrich | embed, cluster, refine |
embeddings and grouping |
| Review | review, remap-labels |
human review and relabeling |
| Export | export {voc,yolo,id,html,stats}, query |
annotations, galleries, numbers |
mbariml review — the curation GUI. Left: the ROI mosaic, tinted by label
with a green check on verified ROIs, the selected tile outlined in blue.
Right: the full source frame with every detection overlaid as a draggable box
(the selected one in red), over the controls panel: adding ROIs by hand or
with SAM3, tightening boxes, sorting and similarity search, display-only
brightness/contrast and color correction, a confidence range, and
verify/relabel/delete.
Also in this repo: cheat_sheet.txt (one worked
example per command, plus verbatim --help for all of them) ·
docs/SCHEMA.md (what's in the database, column by
column) · CHANGELOG.md (what changed, and why).
pip install -e .This installs the mbariml command. Because it's an editable install,
editing any .py file takes effect the next time you run mbariml — no
reinstall needed. Re-run pip install -e . only when pyproject.toml
changes (a new dependency, mainly). (requirements.txt is also kept, fixed,
for anyone who just wants pip install -r requirements.txt — see "What
changed" below for why it needed fixing.)
Optional: SAM3 in the review GUI. mbariml review can use Meta's SAM3
to draw a box from one click and to tighten loose boxes (see "Using the
review GUI well"). Nothing else needs it, and review works exactly the same
without it — the SAM3 buttons are just greyed out, with the reason as their
tooltip. To turn it on, once:
pip install -e ".[sam3]" # adds Ultralytics' CLIP (not on PyPI; PyPI's `clip` is an unrelated tool)
hf auth login # after requesting access at https://huggingface.co/facebook/sam3 (gated)
mbariml sam3 download # sam3.pt (~3.4 GB) into the Hugging Face cache
mbariml review DB_PATH # finds it there; no flag or env var neededmbariml sam3 check says which model review will use, or what's missing.
Review looks for the model in this order: --sam3-model /path/to/sam3.pt,
then $MBARIML_SAM3_MODEL, then the Hugging Face cache. So a sam3.pt you
already have elsewhere works without downloading it again.
It needs a GPU (CUDA or Apple MPS) to be pleasant: on an M3 Ultra it loads in about 6 s on first use, then takes ~0.3–0.45 s per new image and ~15–50 ms per click.
Every command reads/writes the same database schema and just operates on
whatever database you point it at — there's no hidden state, and no command
needs any specific earlier one to have run, only a database that already has
what it needs (embed needs ROI blobs, cluster needs embeddings). So the
Import and Generate commands are all standalone entry points:
mbariml import yolo dataset/images dataset/labels dataset/names.txt /data/results/
mbariml infer images runs/train/best.pt /data/new_survey/ /data/results/
mbariml infer video runs/train/best.pt /data/dive_video/ /data/results/...and because they write the same schema as everything else, you can continue straight into anything downstream:
mbariml review /data/results/yolo_predictions.duckdb
mbariml embed /data/results/yolo_predictions.duckdb
mbariml export html /data/results/yolo_predictions.duckdb /data/results/htmlAll commands:
| Command | Stage | What it does |
|---|---|---|
mbariml import yolo |
Import | Import an existing YOLO dataset (images + labels + names file) (see below) |
mbariml import voc |
Import | Import an existing Pascal VOC dataset (images + XML) (see below) |
mbariml infer images |
Generate | Detect on a directory of images; crop ROIs into a database (see below) |
mbariml infer video |
Generate | Detect on video, by tracking or frame striding (see below) |
mbariml embed |
Enrich | Compute a DINOv3 embedding for every ROI |
mbariml cluster |
Enrich | Cluster embeddings with EVoC, name clusters, export review grids (see below) |
mbariml refine |
Enrich | Re-cluster one label's ROIs into finer sub-clusters |
mbariml review |
Review | Interactive GUI for labeling/deleting/adding ROIs |
mbariml remap-labels |
Review | Bulk-rename new_label values from a CSV |
mbariml export voc |
Export | Pascal VOC XML |
mbariml export yolo |
Export | YOLO label files + names.txt + train/val/test splits (see below) |
mbariml export id |
Export | *.id sidecar next to each source image (see below) |
mbariml export html |
Export | Paginated HTML gallery of images + crops |
mbariml export stats |
Export | Label counts, boxes-per-image stats, image × label matrix (see below) |
mbariml query |
Export | Ad hoc SQL against a database |
mbariml run |
— | Chain ingest → embed → cluster → export |
mbariml sam3 download |
— | Fetch the optional SAM3 model for review (see Setup) |
mbariml sam3 check |
— | Show which SAM3 model review would use, or what's missing |
Run mbariml <command> --help for the full option list (mbariml infer --help / mbariml import --help / mbariml export --help / mbariml sam3 --help for the groups).
mbariml import yolo dataset/images dataset/labels dataset/names.txt /data/results/
mbariml import voc dataset/images dataset/Annotations /data/results/The reverse of export yolo/export voc: no model, just labels someone
already drew. Every box becomes an ordinary row in
/data/results/yolo_predictions.duckdb, ROI crop and sharpness included, so
from there it is exactly like detector output: review can edit, relabel,
delete and add boxes on those images, embed/cluster pick it up, and every
export writes it back out. Point it at an existing database's directory and
it appends, so a legacy training set and a new survey's detections can live
in one database.
- Imported boxes are verified by default. A training set is human-labeled
ground truth, and exports select only verified rows -- imported unverified,
it would export back out as nothing. Pass
--unverifiedwhen the labels are some other model's predictions that still need review. - Pairing labels with images. By relative path first
(
labels/train/a.txt↔images/train/a.jpg), then by bare filename when that is unique. VOC uses each XML's<filename>(with<folder>breaking a tie), never its<path>, which is only right on the machine that wrote it. A label file whose image can't be found, or matches more than one, is skipped and counted. - YOLO names file:
names.txt(one per line) or an Ultralytics dataset.yaml. Every label file is checked against it before anything is written, so the wrong names file fails at once rather than mislabeling every box. A 6th column (save_confconfidence) and segmentation polygons (imported as their bounding box) are accepted too. - Confidence comes from the file where it has one (the YOLO 6th column,
or the
<confidence>thatexport vocwrites), else 1.0, as for a box drawn in review. - Re-running is safe: images already in the database are skipped
(
--no-skip-existingto add their boxes anyway). - YOLO background images (an empty label file, or no label file at all) have no boxes, so there is nothing to store for them; they are counted, not imported.
Then mbariml embed as usual to make the imported ROIs searchable and
clusterable.
infer images is the merge of what used to be two nearly-identical commands,
detect and infer-images (v0.11.0). They wrote the same schema and cropped
ROIs the same way; the entire real difference was batching, annotated-image
saving, and a 16× gap in default confidence — flags and defaults, not
architecture. Keeping two copies meant every fix had to be made twice. The
intent distinction survives as --preset:
--preset curate(default) — conf 0.005, imgsz 1952, no annotated images. Mine everything, then cluster/review and discard the noise. Default deliberately: an over-permissive threshold is recoverable (filter later — the review GUI even has a confidence-range slider), while a too-strict one silently drops detections you can't get back without a full re-run.--preset predict— conf 0.08, imgsz 992, saves annotated images. Believable predictions over new imagery.
Any individual option overrides its preset.
mbariml infer video models/best.pt /data/dive_video/ /data/results/--mode track (default) keeps ONE ROI per tracked object. This is what
you want for building curation/training data from video: a sponge in view for
300 frames is one animal, not 300 training examples — 300 near-identical
crops would swamp clustering and be tedious to review. Measured on five
one-minute benthic clips: 272,126 tracked detections collapsed to 1,163
tracks → 1,163 ROIs.
Tracking is necessarily two passes, because a track's representative frame can't be chosen until the track has ended:
- Track the whole video, recording per-track metadata only (frame, box, confidence, class). No pixels retained, so memory is O(open tracks).
- One forward sweep that extracts just the chosen frames.
Pass 2 sweeps rather than seeks deliberately: CAP_PROP_POS_FRAMES is
unreliable on long-GOP encodings and lands on the nearest keyframe, which
would silently pair a detection's box with the wrong pixels. Decode is far
cheaper than inference, so the extra pass costs a fraction of pass 1.
The tracker has confidence thresholds of its own, and they decide how many
tracks you get, not --conf. A detection starts a track only if it clears
both track_high_thresh and new_track_thresh. Ultralytics' configs set
those at 0.25 (ByteTrack, BoT-SORT) to 0.7 (TrackTrack), far above the
curate preset's 0.005, so on faint benthic footage they throw away nearly
everything: fewer than 0.1% of detections on five one-minute clips reached
TrackTrack's 0.7, and a 10-second clip gave 0 tracks.
So --tracker auto, the default, uses mbariml's own ByteTrack config for the
preset (src/mbariml/trackers/):
| preset | detection --conf |
track starts at | track continues at | lost track kept |
|---|---|---|---|---|
curate |
0.005 | 0.01 | 0.005 | 300 frames |
predict |
0.08 | 0.1 | 0.08 | 300 frames |
The same 10-second clip gives 36 tracks with curate, 18 with predict. On
five one-minute clips curate gave 1,163 tracks, and in a random sample of 24
of the longer ones, every track stayed on one object. (That check was made from
contact sheets by an AI model, Claude, not by a person.)
--tracker also takes an Ultralytics shipped name (tracktrack.yaml,
botsort.yaml, bytetrack.yaml, ocsort.yaml, deepocsort.yaml,
fasttrack.yaml) or a path to your own YAML. If its start threshold is more
than twice --conf, infer video warns that faint objects may produce no
tracks at all.
--track-roi chooses which frame of a track becomes its ROI, and
--track-third which third of the track the -third policies pick from:
best-conf-third(default) — most confident frame of the chosen third.sharpest-third— least blurry frame of the chosen third (Laplacian variance), when crop quality matters more than detector confidence.best-conf/center— whole-track alternatives.--track-third first | middle | last—middleby default.
mbariml infer video models/best.pt dive.mp4 /data/results/ --track-third lastThe middle third is the default from annotators' experience: a track's first
and last frames are when the animal is entering or leaving view — clipped at
the frame edge, occluded, motion-blurred — and plain max-confidence happily
picks exactly those. Measured on 684 benthic tracks (the paper's track-selection analysis),
confidence peaked most often in the last third, but the last third also
had the most boxes touching the frame edge (23%, against 8% in the middle), and
taking the most confident frame anywhere picked an edge-touching box for 42%
of tracks against 9% for the default. Which third gives the better training
example wasn't measured, so it's your choice: --track-third last if your own
footage says so.
best-conf-central and sharpest-central, the names before --track-third,
still work and mean the -third policy with the third you give (middle by
default). Tracks too short for thirds to mean anything use the whole track.
--mode stride skips all of that: sample every Nth frame, treat each as
an independent image, one pass.
Why the frames get written to disk. Each frame that produced a detection
is extracted to a real JPEG under OUTPUT_DIR/frames/, and image_path
points at that file. Video rows are therefore indistinguishable from image
rows to everything downstream — review, embed, cluster, all four exports,
html, stats all work unchanged, with no video-aware code anywhere else in
the pipeline (verified end-to-end). The alternative — storing the video
path plus a frame number and resolving lazily — would have meant teaching
five separate consumers to decode video, and the exports would have had to
materialize frames anyway: a VOC XML pointing at "video.mp4, frame 1234"
isn't something any trainer understands. Frames with no detections are never
written.
The trail back to the footage is kept as provenance columns (video_path,
frame_number, frame_time_s, track_id, track_length) — what the review
GUI's Open Video button uses.
One database or several? Either. Generate commands allocate ids from the
database's own counter, so a second run appends rather than colliding —
several videos, or images and video together, can share one database and be
clustered/reviewed as one set. Verified: 5 image rows + 4 video rows in one
database (ids 0–4 and 5–8), with stats aggregating across both.
cluster names each cluster after the dominant original label of its members.
That works when the detector has enough classes to tell clusters apart, and
fails badly when it doesn't. Run a single-class detector — MBARI's Megalodon,
say, which reports only object — and every cluster's dominant label is the
same string, so writing it back collapses the clustering you just computed into
one undifferentiated label. The grouping survives only in evoc_clust, and the
review GUI (which filters and sorts on new_label) can no longer tell the
groups apart.
--naming controls this, and defaults to auto:
auto(default) — use the dominant label, but when one label wins several clusters, suffix each with an index. A single-class run yieldsobject_1,object_2, … Measured on a real 147-ROI run: 6 clusters that previously all becameobjectnow come back asobject_1–object_6.dominant— always use the bare label (the behaviour before v0.13.0).indexed— always add the index.
auto also helps multi-class runs. On the same ROIs with real labels, one
cluster each of Actiniaria and Ceriantharia keep their bare names while four
separate sponge clusters become Hexactinellida_1–_4, rather than being
flattened into a single Hexactinellida. Merge them later with remap-labels
if that's what you want.
--label-source picks which label the naming vote reads:
original(default) — the raw detector class, i.e. thelabelcolumn.new— your curatednew_label, falling back to the raw class for rows you haven't reviewed yet (the same "curated where curated, raw where not" conventionstatsandexport htmluse).
This matters once a database has been through review, because that's exactly
when the raw class is the useless column. Clustering a reviewed single-class
database with the default names every cluster from object and hands back
object_1, object_2, … — throwing away the taxon names you typed. Measured
on a 300-ROI database curated into three taxa:
resulting new_label |
|
|---|---|
--label-source original (default) |
object_1, object_2, object_3 |
--label-source new |
Muusoctopus, Sponge_sp_A, Coral_bamboo |
With a third of the rows left unreviewed, those rows' cluster falls back to the
raw object while the two reviewed clusters keep their curated names.
Clustering overwrites
new_labelfor every embedded row — verified rows included — whichever source you pick. It does not merge, and there is no undo.verifiedis not cleared either, so afterwards those rows still look human-reviewed while carrying a machine-generated label. Never run it against a curated database you care about — work on a copy.
# 1. Close the review GUI first, then confirm nothing is mid-write.
# A .wal sidecar means an open connection hasn't flushed — the copy
# would be incomplete.
ls /data/results/yolo_predictions.duckdb*
du -h /data/results/yolo_predictions.duckdb # ROI blobs make these big
# 2. Copy into its OWN directory. `cluster` writes roi_grids/ next to the
# database, so a same-directory copy drops grids on the originals.
mkdir -p ~/Desktop/cluster_experiment
cp /data/results/yolo_predictions.duckdb ~/Desktop/cluster_experiment/
# 3. Record what your curated labels look like now, to compare against.
mbariml export stats ~/Desktop/cluster_experiment/yolo_predictions.duckdb
# 4. Cluster the copy, naming the clusters from your curated labels.
mbariml cluster ~/Desktop/cluster_experiment/yolo_predictions.duckdb \
--label-source new --approx-n-clusters 24 --seed 42
# 5. Compare, then look at the actual groupings.
mbariml export stats ~/Desktop/cluster_experiment/yolo_predictions.duckdb
mbariml review ~/Desktop/cluster_experiment/yolo_predictions.duckdbYour real database at /data/results/ is never opened by any of this.
To try different settings, start each attempt from a fresh copy — not by
re-running cluster on the same file:
rm -rf ~/Desktop/cluster_experiment
mkdir -p ~/Desktop/cluster_experiment
cp /data/results/yolo_predictions.duckdb ~/Desktop/cluster_experiment/
mbariml cluster ~/Desktop/cluster_experiment/yolo_predictions.duckdb \
--label-source new --approx-n-clusters 40 --noise-level 0.1 --seed 42That matters because --label-source new reads the very column clustering
writes back to: a second run on the same file votes on the first run's
generated names as if they were human decisions. It's for the first pass over
a reviewed database, not for repeated re-runs.
Always pass --seed while tuning, or EVoC returns a different result every run
and you can't tell whether changing --approx-n-clusters did anything. The copy
carries the embeddings, so embed never re-runs and each iteration takes
seconds — the expensive stage is already done.
If an experiment turns out better, there's no merge path back: re-review in the
copy and adopt it as your working database, or carry specific renames across
with remap-labels. Nothing will splice new clusters into your curated database
while preserving the human labels.
mbariml review DB_PATH labels/deletes ROIs one page at a time:
-
Click a thumbnail to select it, Shift-click to select a range, Ctrl-click to add to the selection.
-
Type a label into the relabel field (free text, or pick from the dropdown of labels already in use) and press Enter, or click Label, to apply it to the current selection.
-
Delete removes the current selection (asks for confirmation first — it's permanent). So does right-click → Delete selected... on a tile that's part of the selection; on a tile that isn't, right-click → Delete... removes just that one (also confirmed first).
-
Escape clears the selection.
-
Confidence range (two-handle slider, then Show): shows only ROIs whose YOLO confidence is in the range, both ends included -- drag the ends to, say, 0.20 and 0.30 to review just that band. Click a handle (or Tab to the slider) and the arrow keys move it by 0.01, Page Up/Down by 0.10. Rows with no confidence show only at the full 0.00-1.00 range. The page counts, the Verified/Unverified counter and "Find similar" all follow the range. View-only: nothing stored changes.
-
Search (type or pick a label, Enter) shows only ROIs whose effective label matches: the new label where one is set, otherwise the original label -- the same label the tile shows in bold. A row renamed from
trashtoanimalis found byanimal, not bytrash. The dropdown lists exactly the labels a search can find;--labelon the command line filters the same way, and "Find similar" ranks within the search by the same rule. -
Right-click a thumbnail → Find similar to rank every ROI in the database by embedding similarity to that one (most similar first), respecting the current
--labelfilter, "Hide verified"/"Hide unverified" and the confidence range if set. This is a whole-dataset re-ranking, not a re-ordering of the page you're on: the closest matches are pulled onto page 1 from wherever in the database they were, and every following page continues down the same ranking. The status line says so explicitly — "sorted by similarity to ROI #812 (all 8,412 matching ROIs)". If it instead reads "3,001 of 8,412 matching ROIs — 5,411 not embedded yet", the search was narrowed by missing embeddings, which is the one thing that can narrow it: runmbariml embedto include the rest. Requiresmbariml embedto have run first. Changing the Sort dropdown exits similarity mode and returns to normal sorting. "Find similar with same label" narrows the ranking to the clicked tile's label, by the same rule everywhere: a tile's name is its new label if it has one, otherwise its original label. So a never-renamedAtile finds every other tile namedA-- renamed toAor stillAfrom the detector -- but not one since renamed toB. -
Sort by "New Label" uses that same rule: a tile with no new label sorts under its original label, in among the renamed tiles, rather than all of them collecting at the end.
-
The status line under the buttons always shows the current page, how many ROIs are shown, and how many are selected, so a keypress never surprises you.
-
Brightness / Contrast sliders (below the tile-size slider) adjust the ROI thumbnails across the whole grid — faint animals against sediment, or low-contrast crops from deep footage, are often far easier to identify stretched than at native exposure. Contrast pivots around mid-grey, so the two sliders act independently rather than fighting each other, and Reset returns both to neutral. View-only: the stored ROI is never modified, so this changes nothing you export.
-
Color-correct display removes the green/blue water cast and restores contrast and the weak reds, on the tiles and the full image alike. It's the same code
mbariml-autolabeluses (copied intombariml/gui/colour.py), measured per whole source image so a crop keeps its colour relative to its surroundings. Measurements are cached in<database>_colour.npznext to the database (the same file autolabel uses). View-only: labels, boxes, embeddings and box-edit re-crops all use the original pixels. -
Hide verified / Hide unverified narrow the grid to what still needs review, or to what's already done. Checking one unchecks the other, since both together would hide everything.
-
Open Video (next to Delete): for ROIs that came from
infer video, opens the source footage at the exact moment that detection was made, from the row'svideo_path+frame_time_s. Tries IINA (macOS), then mpv, then VLC, then ffplay — all four honor a start position — falling back to the default browser with a#t=media fragment.Each is looked up on
PATHand then where its installer actually puts it, since those differ: IINA ships its CLI inside the app bundle, VLC's Windows installer doesn't touchPATH, and ffmpeg on Windows is usually an unzipped folder. Install any one of the four. On Linux and in most ML environmentsffplayis already there as part of ffmpeg, though it's a debug tool with no controls, which is why it's tried last. The browser fallback is a genuine last resort: for a local file the OS generally hands thefile://URL to whatever owns that extension, which ignores the fragment and opens at zero, and no browser plays ProRes. Greyed out for ROIs that came from still images, so the button state itself answers "did this come from video?". -
Add New ROI (green button, top of the controls panel): for a detection YOLO missed entirely. Click it, drag a box on the full-image panel, type a label in the prompt that pops up, and repeat -- it stays armed for drawing as many boxes as you need, on the currently shown image or any other you select next, until you click the button again or press Escape. Each finished box is inserted immediately (confidence fixed at
1.0,verifiedset -- a human drew and named it, there's nothing left to review) and shows up in the detail view right away; the grid picks it up shortly after via the normal page refresh. Its embedding is computed right after, in the background (the status line says so), using the exact same DINOv3 model/preprocessingmbariml embeduses -- directly comparable to every other embedding in the database, so a new box is immediately usable by similarity search/clustering, not stuck withembedding IS NULLuntil someone remembers to runmbariml embed(if that background step ever fails, the ROI itself is still saved andmbariml embedremains a safe fallback -- it only ever fills in rows that are still NULL). An empty/cancelled label prompt discards that box -- nothing is written. While armed, existing boxes step aside -- their handles hide and they can't be dragged -- so a new box can be drawn on top of or inside one (a fish in front of a rock, a part within a whole); they become editable again when you click Done or press Escape. Wheel-zoom and right-click "Delete" keep working throughout. -
Add ROI with SAM3 (next to Add New ROI; needs SAM3, see Setup): click an object on the full-image panel -- inside an existing box is fine -- and SAM3 draws a box around it, shown dashed yellow, with a small popup at the click. SAM3 usually offers up to three nested boxes (a part, the object, the object plus surroundings): Tab / Shift-Tab steps through them, starting on SAM3's highest-scoring one. Type or pick the label (it starts on the last one you used) and press Enter to save it, exactly like a hand-drawn box -- confidence
1.0,verifiedset, embedding computed in the background. Escape, or clicking outside the popup, discards it. It stays on for more objects until you click the button again or press Escape; dragging pans. The two add modes are exclusive. -
Tighten Boxes (SAM3): SAM3 refits every box on the shown image, verified or not. Each prompt is the box plus a click at its centre, and the new box is measured from SAM3's mask, kept within the old box and ignoring stray specks of mask. SAM3's box is then the next prompt, until it stops shrinking (up to 4 prompts per box). You see the new boxes dashed yellow and confirm before anything is written. Only real tightening is offered: boxes SAM3 leaves about the same, or shrinks to a small part of (it found something else there), are left alone and counted in the prompt. Right-click a box and choose Tighten with SAM3 to do just that box (it does show a box SAM3 disagrees on, with a warning). On 60 imported Cyprus litter boxes, tightening left a median 0.45 of the original area. On SeaClear, SAM3's boxes were usually tighter than the hand-drawn ones but occasionally cut off a thin or faint end -- a cable's far end, a fish's tail -- which is why every change is previewed.
-
Any change to a box's geometry recomputes its embedding -- dragging it, tightening it -- along with its crop and sharpness. The stored embedding described the old crop, so it is cleared at once and recomputed in the background; if that ever fails,
mbariml embedfills it in. -
Zooming and panning the full image: the mouse wheel zooms around the pointer, one step per wheel notch; dragging empty image pans. Scrolling nobody meant as zoom is ignored -- a Magic Mouse or trackpad scrolls whenever a finger slides on it, clicks included, and keeps "coasting" after the finger lifts. So the panel ignores momentum scroll, scroll while a button is held or within 0.3 s of a click, and sideways swipes, and no single scroll event zooms more than one notch. Right-drag doesn't zoom either (pyqtgraph's default zoomed 2% per pixel moved, so a Magic Mouse right-click with a slight slide made the zoom jump).
Labeling and deleting no longer rebuild the entire page of thumbnails (the original did, on every single click, which is why review used to feel slow) — only the affected thumbnails' captions update immediately. Sorting and paging still do a full rebuild, since the visible set of ROIs actually changes then.
Sort by Sharpness now actually means something. It used to sort by a
column that was hardcoded to 0.0 for every row -- a no-op. mbariml infer images/infer video compute a real blur score (Laplacian variance)
per ROI as they extract it, so every database built since v0.7.0 has real
values to sort by. (A backfill-sharpness utility used to exist for
databases built before that change, recomputing the score after the fact
from each row's already-stored ROI crop; it was removed in v0.9.0 once every
current write path covered it at the source. If you're still holding a
pre-v0.7.0 database that needs backfilling, ask for that utility back rather
than assuming a migration path still exists.)
Sorting by sharpness (ascending) then surfaces the blurriest ROIs first — useful for quickly finding and deleting unusable crops.
mbariml embed uses DINOv3 (ViT-Large/16, vit_large_patch16_dinov3.lvd1689m
general-purpose weights) — swapped in from DINOv2 for accuracy. Embeddings
from different backbones are not comparable: if you ever change
EMBEDDING_MODEL_NAME in src/mbariml/steps/step2_embed.py, or a database
already has embeddings from a different model, re-embed the whole thing with
--force so nothing ends up mixing two different embedding spaces (which
would silently corrupt both clustering and the GUI's similarity sort):
mbariml embed /data/survey_results/yolo_predictions.duckdb --forceEmbedding is batched (--batch-size, default 32) — ROIs are decoded and
run through the model together in one forward pass per batch, instead of one
at a time. One-at-a-time was the original approach and made a GPU/MPS sit
mostly idle waiting on serialized dispatches instead of doing throughput
work. Increase --batch-size for more throughput up to your device's
memory limit; decrease it if you hit an out-of-memory error.
Decoding/preprocessing is parallelized (--decode-workers, default
min(16, cpu_count)) — this CPU-bound work (JPEG decode, resize, normalize)
used to run in a single-threaded Python loop, pinning one core while dozens
sat idle on a many-core machine and leaving the GPU waiting on it.
Database writes are batched (--flush-size, default 2000) — this was
the dominant real bottleneck, much bigger than either point above. New
embeddings used to be written with one UPDATE ... WHERE id = ? per batch.
Measured directly on real hardware: model throughput was a rock-stable ~80
items/sec, but per-batch DB write time grew from ~1s to ~15s over just 45
batches and kept climbing — DuckDB is a columnar/OLAP engine and is
documented to be
dramatically slower at many small row-by-row UPDATEs (each carries MVCC
row-versioning overhead) than at one bulk UPDATE. New embeddings are now
staged into a temp table and applied with a single UPDATE ... FROM every
--flush-size rows instead of once per batch. Verified end-to-end: a run
that degraded from 45 it/s to 6.5 it/s (and was still falling) over 3000
ROIs became a flat ~70-75 it/s for the same 3000 ROIs after this fix — no
degradation at all, ~6x faster overall on top of removing an actively
worsening trend that would have made a large run take dramatically longer
than a naive per-item estimate suggests.
It's safe to interrupt and resume any time: mbariml embed only processes
rows with embedding IS NULL, so re-running after a Ctrl-C (or a crash)
picks up right where the last completed flush left off.
If a long embedding run is still far slower than expected after all of the
above, with a confirmed MPS/CUDA device, rule out something else competing
for the GPU or unified memory (check Activity Monitor / nvidia-smi while
it runs) before assuming it's this code.
On MLX: mlx-image does
have a DINOv3 implementation for Apple Silicon (weights converted from the
same timm/HuggingFace source, hosted at
huggingface.co/mlx-vision), so it's a
real option if you want to explore it further. It wasn't adopted here: no
published benchmark showed it meaningfully outperforming a well-optimized
PyTorch/MPS pipeline for this kind of batched-inference workload (MLX's
biggest advantages are for autoregressive/LLM-style workloads, not a single
forward pass per batch), and — as the numbers above show — the actual
bottleneck was never the model or the framework at all. Porting to MLX would
not have fixed a DuckDB write pattern.
mbariml cluster used to be able to take hours on a large database, with
no visible progress. Directly measured cause: not the clustering math.
evoc.EVoC.fit_predict() isn't brute-force KNN -- it has its own
JIT-compiled approximate nearest-neighbor search built in already (similar
in spirit to DuckDB's VSS/HNSW extension), and clusters 50,000 embeddings in
~2.6 seconds, scaling roughly linearly (300k+ points finishes in well under
a minute). Adding an ANN index wouldn't touch this cost at all -- there's no
repeated similarity-search query here for an index to accelerate, just one
batch clustering call that's already fast.
The actual cost was the same DuckDB small-UPDATE problem as embed (see
above), except worse: writing cluster results touches the indexed
new_label column. Measured directly: 100,000 rows via one UPDATE per
row took 38.6 seconds and was still trending worse; the same 100,000 rows
via a staged bulk UPDATE ... FROM (mbariml.db.bulk_update) took 6.0
seconds. cluster now uses this, logs progress at each phase (fetching,
building the embedding matrix, clustering, writing results) so a long run
is never silent, and step 4 (refine) got the identical fix for the same
reason.
A second, even more fundamental finding while chasing this down:
DuckDB's Python driver commits (and fsyncs to disk) after every individual
statement by default when writing to a file-backed database -- even
within a single executemany() call. Measured directly: an identical
2000-row INSERT took 11.46 seconds without an explicit transaction
around it, and 0.82 seconds wrapped in one (BEGIN TRANSACTION /
COMMIT) -- a ~14x difference from transaction-wrapping alone, independent
of row count or which column is touched. This affected every step that
writes many rows directly to the persistent predictions table:
mbariml.db.fast_executemany (a drop-in replacement for
conn.executemany) now wraps every such call in steps 1, 2, 3, 4, 5, and 8
in an explicit transaction. Steps 1 and 8 already committed incrementally
per image/batch for resumability -- that's unchanged; this fix is about
what happens inside each of those commits, not how often they happen.
Every downstream consumer — export yolo, export voc, export id,
export html and stats — selects and names localizations by one shared
rule, defined once in mbariml.db (EFFECTIVE_LABEL_SQL / curated_where()):
| In the database | Exported? | Name used |
|---|---|---|
| verified, name unchanged | yes | the original detector label |
| verified, name updated | yes | new_label |
| not verified | no | — |
The reason this needs stating: the review GUI's Verify button sets
verified = 1 without writing new_label — only relabelling writes it.
So new_label IS NOT NULL means "boxes whose name I retyped", not "boxes I
confirmed", and the three dataset exports used to filter on exactly that.
On a 35,492-row survey database that exported 2,805 boxes and silently
dropped 32,687 confirmed ones; worse, 29,917 of the dropped boxes sat on
images that were in the export, so YOLO read them as unlabeled background
and trained against the reviewer's own identifications.
export html and stats are diagnostics as well as previews, so they
accept --include-unverified to fall back to summarizing raw detector
output on a database that has not been reviewed yet. export id and
export yolo accept it too, opt-in -- yolo for bootstrapping a dataset
from model predictions (still excluding noise); verified-only stays the
default everywhere. export voc is the only export without it.
cluster is the deliberate exception: it groups and relabels unverified
data too, since finding names for un-reviewed ROIs is the whole point of it.
Every export that writes one file per source image — export yolo's
labels/, export voc's XML, export id's sidecars, export html's
images and crops, the image manifest, and the stats matrix rows — names
that file after the source image's own filename:
/Volumes/.../PROSILICA_R/1784394530578235.tif
-> labels/1784394530578235.txt
-> pascal_voc/1784394530578235.xml
-> 1784394530578235.id
There is one exception, and it matters. Flattening a nested mission tree
into a single output directory is exactly when dive01/img_0001.jpg and
dive02/img_0001.jpg become the same output file, and the second silently
overwrites the first — losing a whole image's annotations with no error.
So naming is decided across the export as a whole: if two source images
share a filename, only those fall back to <parent_dir>_<stem>, and the
clash is logged:
WARNING 1 filename(s) appear on more than one source image in this export, so
those output files keep the <parent_dir>_<filename> prefix to avoid
overwriting each other: img_0001 (2 images)
One unlucky pair in a 10,000-image survey therefore doesn't prefix the other
9,998. The same map names every artifact of a given export, so labels/,
the split lists and the manifest can never disagree about what an image is
called. If two images somehow resolve to the same name even after
disambiguation, the export fails rather than overwriting.
mbariml export yolo DB_PATH OUTPUT_DIR writes a complete, trainable
dataset skeleton from curated labels — every verified localization,
named new_label where you retyped it and the original detector label
where you confirmed it unchanged, excluding noise. See
What counts as a curated label:
| Output | What it is |
|---|---|
labels/<name>.txt |
one per image, each line class_id x_center y_center width height normalized against that image's actual pixel dimensions |
names.txt |
class index → label, in the order the label files use |
train.txt / val.txt / test.txt |
the image lists training configs point at — one ./images/<file> path per line |
<dataset>.yaml |
the Ultralytics dataset config — split paths, nc, names (see below) |
image_manifest.csv + copy_images.py |
fetch the matching images (see below) |
mbariml export yolo /data/survey_results/yolo_predictions.duckdb /data/survey_results/yolo_out/
python3 /data/survey_results/yolo_out/copy_images.py --dest /data/survey_results/yolo_out/imagesThose two commands leave a directory that trains as-is.
Blending into an existing dataset. Class indices are assigned
alphabetically by default, so a separate export's class 3 need not be the
existing dataset's class 3. Pass --names-file with that dataset's
names.txt or dataset .yaml to keep its exact order: every name keeps its
index, including ones this export doesn't use, and any label it doesn't list
is appended at the end (logged, with a hint when it differs from an existing
name only by case). The written names.txt/YAML is then the full, extended
list — use it for the blended dataset.
mbariml export yolo new.duckdb new_export/ --names-file existing_dataset/existing.yamlExporting only confident boxes. --conf (default all) sets a minimum
detection confidence, as a percentage (50, 70, 85%) or a fraction
(0.7). It filters whole images: an image is exported only if every box
it would export meets the threshold, because dropping just the low box would
leave that object in the image unlabeled, which YOLO learns as background.
A verified box counts as 100% whatever the detector scored it (a human
confirmed it; hand-drawn boxes are verified too), so the threshold only ever
applies to unverified boxes -- use it with --include-unverified, to take
the verified data plus only the confident part of the raw model output. On
its own it has no effect (and says so). The log reports how many images and
boxes were dropped.
mbariml export yolo predictions.duckdb dataset/ --include-unverified --conf 70The dataset YAML. Named <output_dir name>.yaml by default
(--yaml-name to change it), in the format Ultralytics expects:
# train and val data
train: train.txt
val: val.txt
# number of classes
nc: 51
# class names
names: ['Actiniaria',
'Actinopterygii',
...
'tube']nc and names are written from the same in-memory list that assigned the
class indices in labels/ and produced names.txt, rather than being
recomputed from the database — so the three cannot disagree about which
index is which taxon. That matters more than it sounds: a YAML whose name
order differs from the indices in the label files trains every class against
the wrong name and looks completely normal while doing it.
The split paths are relative, and there is no path: key. Ultralytics
resolves them against the YAML's own directory when path is absent, so
naming the YAML when you kick off training is all that's needed — the same
dataset directory works read from /Volumes/M3_ML/... on a Mac or
/mnt/M3_ML/... on the Linux trainer, with nothing to rewrite in between.
(Verified against Ultralytics 8.4.154, including after moving the folder.)
An empty split is left out of the YAML entirely. With --split-ratios '85 15 0' there is no test set, so no test: key is written — rather than
one pointing at an empty test.txt, which would fail later, at evaluation
time, long after the export looked fine.
A typical export, 85/15 train/val with no test split:
mbariml export yolo /data/survey_results/yolo_predictions.duckdb \
/Volumes/M3_ML/training_data/2026/MBARI_lassml_my_survey_20260916/ \
--split-ratios '85 15 0' \
--yaml-name MBARI_lassml_my_survey_20260916_yolo26s_LL.yamlPass --no-yaml to skip it. It needs --splits (which is the default),
since it points at the split files.
The split paths are relative on purpose. ./images/<file> means the
dataset directory only ever refers to itself, so it can be zipped, copied to
a training box, or moved between volumes without a single path needing to be
rewritten — which absolute paths recorded on the machine that ran the export
could not survive. Every filename in a split file is the same name
copy_images.py copies to and labels/ is keyed by — the source image's own
filename (see Output filenames) — so images/X.jpg ↔
labels/X.txt pairs up by construction — exactly the
pairing YOLO resolves by swapping /images/ for /labels/ in these paths.
| Option | Default | Notes |
|---|---|---|
--split-ratios |
"85 10 5" |
train/val/test percentages; must sum to 100, and a set that doesn't is rejected rather than rescaled (a typo'd "80 10 5" means a miscount, not a request to drop 5% of the data on the floor) |
--split-seed |
42 |
same database + ratios + seed always reproduces the same split |
--test-images-file |
— | a file of image filenames (one per line) to pin into test every time — a fixed benchmark set held out across every export. Matched leniently: exported name, original filename, full path, with or without extension |
--no-splits |
— | skip the split files entirely |
Seeded by default deliberately: re-exporting after relabelling a handful of ROIs should not silently reshuffle which images were held out, or every model trained before and after the re-export becomes incomparable.
Splitting happens per image, never per box. Two crops of the same frame landing on opposite sides of the train/val boundary leak the identical background, lighting, and often the same individual animal across the split — which quietly inflates validation scores on benthic transect imagery, where consecutive frames already overlap heavily.
Images recorded in the database but no longer on disk are left out of the splits (and the manifest) rather than listed: a split line pointing at an image that was never copied surfaces much later as a training-time error.
It deliberately doesn't copy the source images into an images/ folder
itself (they already exist on the survey volume this ran against, and
copying every JPEG would duplicate the lot). Instead,
export yolo and export voc both also write image_manifest.csv (every
distinct source image referenced, mapped to its destination
filename) and a standalone copy_images.py next to it. Run that script
later — from this machine or any other that can see the recorded source
paths — to actually pull the matching images down. The voc script
defaults to Desktop; the yolo script requires --dest (alias
--output-dir), since its images belong in the images/ directory next to
the label files:
python3 /data/survey_results/voc_out/copy_images.py
python3 /data/survey_results/yolo_out/copy_images.py --dest /data/survey_results/yolo_out/imagescopy_images.py is stdlib-only (argparse/csv/shutil/pathlib) and
doesn't import mbariml — it's meant to be portable, not tied to this repo
being installed wherever it eventually runs.
mbariml export id DB_PATH writes a <image_stem>.id sidecar file next to
every source image that has at least one curated identification (every
verified localization, named new_label where you retyped it and the
original label where you confirmed it unchanged, excluding noise) —
wherever that image actually lives on disk, so it
naturally follows a nested mission directory structure:
mbariml export id /data/survey_results/yolo_predictions.duckdb--output-dir collects them into one directory instead, for a read-only
survey volume or a handoff that doesn't include the imagery:
mbariml export id /data/survey_results/yolo_predictions.duckdb --output-dir ~/Desktop/idsFiles there are named after the source image (<stem>.id), with the
parent-directory prefix added only where two images would otherwise collide —
see Output filenames. Each file's header records the
full source path regardless.
Each file has a commented header — generator + version, who ran the export,
the model that produced the detections (recorded automatically by the infer
command, or override with --model), the full path of the source image,
its pixel dimensions, the identification count, and a legend for the columns
— followed by one CSV row per identification:
# mbariml identification file
# generator: mbariml v0.20.0
# generated_by: lonny
# generated_at: 2026-09-17T21:55:59Z
# model: /path/to/best.pt
# source_image: /Volumes/SeafloorMapping/2026/20260718d1/images/.../1619554491865857.png
# image_width: 1936
# image_height: 1456
# count: 2
#
# ... field legend ...
#
# index,label,confidence,center_x,center_y,lon,lat,depth,tl_x,tl_y,tr_x,tr_y,br_x,br_y,bl_x,bl_y
0,Crinoidea,0.9463,1499,482,0.0,0.0,0.0,1470,454,1528,454,1528,510,1470,510
1,marine organism,0.6110,67,230,0.0,0.0,0.0,40,200,95,200,95,260,40,260
Every value is comma-separated — the rows are plain CSV and the last
comment line names the columns, so csv.reader (or pandas, or a spreadsheet)
reads them directly. Rows are written with csv.writer, which quotes only
when it must: none of this survey's 51 labels contains a comma, but one that
did would come out as "Nudibranchia, sp. A" rather than silently adding a
column and shifting every coordinate after it.
lon,lat,depth appear once, next to the center, because they describe
where the observation is and an observation has one position — a separate
navigation-merge process fills them in per center point. They used to be
repeated per vertex, so every row shipped five identical 0.0,0.0,0.0
triples for a single unknown.
Those three are the only columns that process should change (0-based 5, 6, 7); everything else is written back unchanged, with a CSV-aware writer so a label that needed quoting stays quoted. Filling them in is then just:
for row in csv.reader(data_rows):
row[5], row[6], row[7] = lon, lat, depth # fix for (center_x, center_y)
writer.writerow(row)center_x/center_y are the observation's position — one point to put
on a map or match to a navigation fix — and the corners give the extent. The
center is the exact integer midpoint of the tl/br corners beside it, not
a separately-rounded midpoint of the underlying floats: those differ by a
pixel on 222 of this survey's 35,492 rows, and a file whose stated center
disagrees with its own corners is the confusion the field exists to remove.
Corner coordinates are box edges, not pixel indices. A box flush
against the right side of a 1936-wide frame has tr_x 1936 — one past the
last column (1935). That is correct, not an off-by-one: it is what makes
width = tr_x - tl_x exact, the same arithmetic export yolo uses. On this
survey 422 boxes touch the right edge and 13 the bottom, so clamping them
would quietly shrink 435 boxes by a pixel. center_x/center_y are true
pixel indices and always stay within 0..W-1 / 0..H-1.
image_width/image_height record the frame size, so a consumer can
bound-check a coordinate without opening the imagery. They come from the
image header only (PIL's lazy open), not a decode — 6.3 ms per image rather
than 43.9 ms, about 6s instead of 44s across a 995-image export. An image
that has moved writes unknown; the identifications are unaffected, since
all box geometry comes from the database.
source_image is the full recorded path for the same reason it matters
under --output-dir: the basename alone doesn't say which dive an
identification came from, and a survey holds many directories with
same-named images.
mbariml export stats DB_PATH prints two tables to the console: label counts (with
percent of total) and boxes-per-image summary stats (avg/min/median/max,
across every image with at least one detection). Both count the same
population the exports write — verified localizations, named
COALESCE(new_label, label) — so these numbers are a reliable preview of
what a training set will contain. Pass --include-unverified to count raw
un-reviewed detections too, which is what makes this useful on a database
fresh out of infer, before any review. noise is included
by default (useful while curating, to see how much of the database is still
noise/unlabeled); pass --exclude-noise once you want real-identification
counts only:
mbariml export stats /data/survey_results/yolo_predictions.duckdb
mbariml export stats /data/survey_results/yolo_predictions.duckdb --exclude-noise --top 20Pass --output-dir to also write label_by_image_matrix.csv — an image ×
label count matrix (rows are images, columns are labels, values are box
counts) for granular, per-concept-per-image analysis, e.g. loading straight
into pandas/R for ecological statistics (per-image richness, per-label
frequency-of-occurrence across images, etc.):
mbariml export stats /data/survey_results/yolo_predictions.duckdb --output-dir /data/survey_results/stats/mbariml run best.pt /data/survey_images/ /data/results/This chains ingest → embed → cluster → export (voc + html) against
/data/results/yolo_predictions.duckdb. Video works too — the chain is
identical after ingest:
mbariml run best.pt /data/dive_video/ /data/results/ --media videoUse --from/--to with named steps (ingest, embed, cluster,
export) to run only part of it — e.g. to resume after reviewing in the GUI
and just re-export:
mbariml run best.pt /data/survey_images/ /data/results/ --from exportThe export step runs both export voc and export html; if you only want
one, run it directly rather than through run. review, query,
remap-labels, export stats, export yolo, and export id aren't in the chain —
they're interactive, or they don't belong in the middle of a batch run.
| Symptom | Cause / fix |
|---|---|
infer finds no detections at all |
Check the model path resolves, then the threshold: --preset curate uses conf 0.005, --preset predict uses 0.08. A model trained on different imagery may genuinely find nothing. |
infer video --mode track finds few or no tracks |
Track creation is gated by the tracker's thresholds, not --conf. Use --tracker auto (the default), whose thresholds match the preset. With your own --tracker, lower its track_high_thresh / new_track_thresh; a warning at startup says when they are far above --conf. Lowering --conf alone will not help. |
| An export reports "N image(s) could not be found on disk" | The database references images that have moved, or a volume that isn't mounted. Paths are recorded when rows are added (absolute since v0.11.0); re-run infer or import if the imagery has been relocated. |
cluster says "too few to cluster" |
EVoC needs more rows than --n-neighbors (default 40). Lower --n-neighbors, or drop --limit. |
| Right-click similarity sort says "no embedding" | Run mbariml embed on the database first. |
| Similarity sort looks like it only sorted the current page | It never does — it ranks the whole matching set. Check the status line: it reports the pool as "all N matching ROIs", or "M of N — … not embedded yet" when a partial/interrupted embed is the limit. A --label filter, "Hide verified" or a confidence range also narrow the pool by design. |
| Clustering or similarity results look nonsensical | Check you haven't mixed embeddings from two models in one database. If you changed EMBEDDING_MODEL_NAME, re-embed everything with mbariml embed --force. |
embed is slow, and getting slower |
Confirm the device (it logs MPS/CUDA/CPU at startup), then check nothing else is competing for the GPU. The historical cause was a DuckDB write pattern, long since fixed — see CHANGELOG.md. |
| "Open Video" does nothing useful | Install IINA (macOS), mpv, VLC, or ffmpeg/ffplay — all honor a start position. Without one it falls back to your browser, which for a local file usually just hands off to the default app and opens at zero. |
| "Open Video" reports success but no window appears | Fixed in v0.21.1. iina-cli could misdetect standard input and pass --stdin to IINA, which then waited for media on stdin and never played the file. If you are on an older version, upgrade or launch the player by hand. |
| Opening an older database errors on a missing column | Only review and the two infer commands open through init_curation_db, which runs the ALTER TABLE ... ADD COLUMN IF NOT EXISTS migrations; the rest open the file as-is. Open it once with mbariml review to migrate it in place, then re-run whatever failed. |
The design, the workflow, and what has and has not been measured are written up in:
Lundsten, L., Barnard, K., & Caress, D. (2026). mbariml: a curation pipeline for turning deep-sea imagery and video into object-detection training data. arXiv:2609.25500. https://doi.org/10.48550/arXiv.2609.25500
@article{lundsten2026mbariml,
title = {mbariml: a curation pipeline for turning deep-sea imagery and
video into object-detection training data},
author = {Lundsten, Lonny and Barnard, Kevin and Caress, Dave},
journal = {arXiv preprint arXiv:2609.25500},
year = {2026},
doi = {10.48550/arXiv.2609.25500},
url = {https://arxiv.org/abs/2609.25500}
}GitHub's "Cite this repository" button reads
CITATION.cff, which names the paper as the preferred citation
and the software as the fallback.
MIT — see LICENSE. Copyright (c) 2026 Monterey Bay Aquarium
Research Institute (MBARI).
Parts of the review GUI are adapted from MBARI's
vars-gridview, also MIT. See
THIRD_PARTY_NOTICES.md for that attribution and for
the licences of the dependencies, which are not all permissive — in particular
Ultralytics YOLO is AGPL-3.0.
Version-by-version history — including the measured performance findings (DuckDB's per-statement fsync, the embedding throughput collapse, the clustering write pattern) and the correctness bugs behind the current design — now lives in CHANGELOG.md.
