Skip to content

Worktree fix npt chunk boundary state - #90

Closed
miketynes wants to merge 11 commits into
mainfrom
worktree-fix-npt-chunk-boundary-state
Closed

miketynes wants to merge 11 commits into
mainfrom
worktree-fix-npt-chunk-boundary-state

Conversation

@miketynes

Copy link
Copy Markdown
Collaborator

No description provided.

…exceptions

An unhandled exception during dynamics execution (e.g. numpy.linalg.LinAlgError
from ASE's MTK-NPT barostat) previously propagated out of the DynamicsRunner's
run loop and crashed the agent, silently stalling the trajectory while the
Slurm job kept reporting RUNNING. Now such failures are caught, logged, and
recorded via mark_trajectory_failed before a clean agent shutdown.
MAD-vs-timestep scatter now colors points by which model_version
produced each frame (via a helper that mirrors get_trajectory_atoms's
chunk iteration/dedup exactly). New "UQ vs error" section plots
DBTrainingFrame.calibration_uq against calibration_error (the same
pair the Controller uses for real threshold calibration), log-log,
colored by model_version_sampled_from.
model_version is a small integer category, not a continuous quantity, so
a viridis colorbar was misleading. Both the MAD plot and the UQ-vs-error
plot now scatter one group per model_version with a proper legend.
Pulls the first-ever frame and the last-ever recorded frame directly
(not via get_trajectory_atoms, which only returns passed chunks) so a
trajectory stuck on a failed/crashed final attempt still gets a real
end-state comparison instead of silently falling back to its last
passed chunk.
No per-member bootstrap provenance is persisted (Trainer.train_model
draws indices with an unseeded RNG, held only in memory), so this
characterizes the round-1 training pool directly instead: every one of
its 200 frames has a labeled-atoms min interatomic distance under 1 A
(98.5% under 0.5 A), i.e. essentially collided structures. A Monte
Carlo simulation of the actual bootstrap procedure on that real pool
shows meaningfully uneven per-member exposure to the worst frames
arises from resampling variance alone, even though every member draws
from the same uniformly pathological pool.
Every chunk of a trajectory ran as a separate executor task, and
advance_dynamics constructed a brand-new ASE NPT/MTKNPT integrator each time.
ASE zeroes the barostat/thermostat extended-system state (cell/barostat
momenta, Nose-Hoover chain variables, or eta/zeta for classic NPT)
unconditionally in __init__, and cascade never persisted or restored it, so
the barostat/thermostat effectively restarted from rest at every chunk
boundary even though atomic positions/momenta/cell carried over correctly.
This produced a visible discontinuity in atomic motion rate at each chunk
seam (e.g. every 1000 steps), independent of any model-version change.

DynamicsRunner now carries the integrator's extended-system state across
chunks (mirroring how self.atoms already carries over, including only
advancing on a passed audit so a retried attempt reuses the same starting
state), threaded through AdvanceSpec -> advance_dynamics -> back to the
runner via new extract_dyn_state/restore_dyn_state helpers in traj_config.py.
@miketynes miketynes closed this Aug 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant