Repository navigation
Conversation
…. A revision One documented module, koopmanrl_utils/skvi_policy_checks.py, trains SKVI with the tuned configurations in configurations/ and reads its deployed Gibbs policy off the Koopman tensor: on the linear system it compares the learned quadratic form and the mean-action gain with the discounted LQR; on the double well it prunes the value function and compares returns on paired rollouts and the total-variation distance to the full policy. Results are written to an untracked output directory; --summarize_only prints the statistics quoted in the ESM and --plot draws its two figures. No stored outputs and no new environment file: it runs in the package's uv environment (uv run -m koopmanrl_utils.skvi_policy_checks). tests/test_skvi_policy_checks.py checks the batched policy against the package's pis and the cost against vectorized_cost_fn. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
22bdafa to
666e7dc
Compare
… SKVI policy checks Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…llects tests that import koopmanrl_utils Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…icy, sensitivity of SKVI) - koopmanrl_utils/skvi_sensitivity_checks.py: --study accuracy (Table S6: one-step error of the tensor along the trained SKVI policy and on random-agent data at a common scale, persistence baseline, state error, exact double-well conditional mean, local one-step Jacobians, closed-loop distance) and --study sensitivity (Tables S7 and S8: dictionary order, action cost weighted by the time step, re-identification along the policy, against no control and the LQR controller of the linearisation). Results carry their settings; the summary refuses to mix runs made with different settings and prints explicit failure counts. - koopmanrl_utils/skvi_policy_checks.py: train_skvi takes optional state-dictionary order, action-cost scale and transitions appended to the identification data (identify_tensor), and quadratic_cost an optional action-cost scale. Defaults unchanged; tensor and value weights bitwise identical on seed 100 for both S11 environments. CONFIG_FILES covers the fluid flow and Lorenz. - tests/test_skvi_sensitivity_checks.py: sampler vs the package's pis, prediction vs K_, double-well conditional mean through the environment's own noise, error metrics, episode cost sums and near-target fraction on a deterministic system, appended identification data, local Jacobian, LQR reference, action-cost scale. - AGENTS.md lists both check modules; .gitignore ignores both output directories. Consolidates the exploratory scripts that produced the S16 tables; on seed 100 it reproduces all 203 reported numbers exactly (the model Jacobians of the sensitivity study now use double precision, differences below 1e-7). Reviewed by Codex (GPT, xhigh); its seven requests are addressed here. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@Pdbz199 @ludgerpaehler could you both take a look at the new commit (34c6d01)? It adds
The updated description has the details and the verification. We won't get this merged before the resubmission, so the data accessibility statement will point to this PR and commit. Thanks! |
…er 0.1% of the episode The summary counted only seeds with no step near the target at all. The ESM states the threshold instead: with the dictionary of order five, three seeds were near the target for 0.00 to 0.01% of the episode, and with order six one seed for 0.02%. A full rerun of all eight seeds reproduces the saved S16 results exactly (1,624 numbers). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@Pdbz199 @ludgerpaehler a quick bump in case this got buried. There are two new commits since Preston's approval: 34c6d01 adds the S16 checks ( A full run of the module on all eight seeds reproduces every number in the S16 tables exactly. The manuscript's data accessibility statement now cites this PR at 12a9edf, so it would help to hear soon if anything needs changing. A short look is enough; the description lists what to check. |
…Preston's review, point 4) - train_skvi takes extra_transitions_for: "both" (default, re-identification as before), "tensor" (the value function is fitted on the random-agent states only) or "value" (random-agent tensor, policy states added to the value fit). SKVI draws its value-fitting batches from dynamics_model.X / Phi_X, so re-identification changes both. - Two Lorenz variants, "order 3, refit x1, tensor only" and "order 3, refit x1, value states only". On eight seeds the median cost is 241,000 tuned, 48,000 re-identified, 97,000 with the tensor refitted only (better on six of eight seeds) and 844,000 with the policy states in the value fit only: most of the gain comes from the tensor. - --variants runs a subset (each variant is seeded on its own, so the numbers equal those of a full run); unknown names are rejected, and a run never replaces a result file made with other settings. - Tests for both guards; 14 tests pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Following Ludger's suggestion, the comments of both modules now describe the computation in full instead of pointing at sections, tables or equations of the supplementary material: - the normalised one-step error is written out (was "equation (4.7)"), with the conditional-mean target and the two choices of scale; - the action-cost scaling says what R * dt does (was "the R * dt weighting of Table S7"); - the single-precision rounding of appended transitions and the generator order in gain_near_target say why they are kept (was "as in the runs reported in the ESM"); - the functions, arguments and constants describe their study instead of naming a table, and the summary prints "Accuracy" and "Sensitivity" instead of "Table S6" to "Table S8". Where the code implements a theoretical statement, the docstring gives the formula and cites the statement by number and name: action_fields (Proposition S11.1, the policy is an exponential family over the action dictionary) and policy_reading (Corollary S11.2, the Gaussian closed form). Each module docstring ends with one reference to the paper's supplementary material, by section title. No computation changes: with docstrings removed, the syntax tree of skvi_policy_checks.py is identical and that of skvi_sensitivity_checks.py differs only in the two printed labels. 14 tests and the pre-commit hooks pass. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
@ludgerpaehler following your point about paper references in code comments, 1cebcf5 rewrites the comments of both modules so that each one describes its computation in full:
Following the middle path you suggested, the two places where the code implements a theoretical statement cite it next to the formula, by number and name: No computation changed. With docstrings removed, the syntax trees are identical apart from the two printed labels, and the 14 tests and the pre-commit hooks pass. Once you have tagged the release with a Zenodo DOI, we can cite that in the data accessibility statement instead of a commit hash. |
koopmanrl_utils/koopman_prediction_validation.py produces the numbers behind the Koopman-model validation of the RSPA revision: the one-step error table and the error-by-degree figure of the main text, and the two figures of the supplementary section "Koopman model validation: additional results" (multi-step error, training against held-out error). For each benchmark the tensor is identified from random-agent trajectories with the tuned SKVI dictionary orders and identification budget, on 80% of the trajectories, and evaluated on the other 20% (split by whole trajectory), over seeds 123-130. The one-step error is measured in dictionary space, the only way SKVI and SAKC use the tensor, against the conditional mean, with the persistence baseline and, for the double well, the noise floor. With the default arguments the module reproduces the files used for the paper: the CSVs agree to floating-point rounding (largest difference 1e-12) and the table rows for Lorenz, fluid flow and double well are identical. The linear system's errors are at machine precision (about 1e-15), so their leading digits change from run to run. Also ignores the output directory, lists the module in koopmanrl_utils/AGENTS.md and adds six tests (3 s), including one that the settings in the module match configurations/. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01XFsKdJUnxiu7MRwnV9jbVZ
Adds the checks behind two sections of the electronic supplementary material (ESM) of the revised manuscript (RSPA-2026-0421): Section S11 ("Reading the policy off the representation", and the closing sentences of "A Note on Interpretability") and Section S16 ("Accuracy along the learned policy and sensitivity of SKVI").
The first three commits are the S11 checks Preston reviewed and approved. The next three add S16, the third of them answering Preston's review of the S16 commits. The newest (1cebcf5) makes the comments of both modules self-contained, following Ludger's suggestion (see Code comments below).
Contents
koopmanrl_utils/skvi_policy_checks.py(S11): trains SKVI with the tuned configurations inconfigurations/and reads its deployed Gibbs policy off the Koopman tensor.train_skvitakes optional arguments (state-dictionary order, action-cost scale, transitions appended to the identification data, and whether they go to the tensor, the value fit or both) andquadratic_costan optional action-cost scale. The defaults are unchanged, and the trained tensor and value weights are bitwise identical to before (checked on seed 100 for both environments).koopmanrl_utils/skvi_sensitivity_checks.py(S16, new):--study accuracy(Table S6, all four benchmarks): one-step error of the tensor identified from random-agent data, along the trained SKVI policy and on fresh random-agent data, with the scale of each dictionary function fixed from the random-agent targets for both distributions. Also the persistence baseline, the error of the predicted state in state units, the exact conditional mean for the double well, the one-step Jacobian of the fitted model at the target against the environment's one-step map, and the distance of the closed loop and of the uncontrolled system from the target.--study sensitivity(Tables S7 and S8, double well and Lorenz): SKVI retrained with other dictionary orders, with the action cost weighted by the time step, and after re-identifying the tensor on transitions of the trained policy, compared with no control and with the LQR controller of the linearisation on the same episodes. Only SKVI is retrained. Two Lorenz variants separate the two effects of re-identification (refit the tensor only, or add the policy's states to the value fit only), and--variantsruns a subset with the same numbers as a full run.--summarize_onlyprints the medians quoted in Tables S6 to S8, including explicit failure counts. Results go to an untrackedskvi_sensitivity_checks_results/, with the settings of each run. The summary refuses to mix runs made with different settings, a run never replaces a result file made with other settings, and unknown variant names are rejected.tests/test_skvi_policy_checks.py(S11) andtests/test_skvi_sensitivity_checks.py(S16): the sampler equals the package'spis, the prediction equalsK_, the double-well conditional mean checked through the environment's own noise, the error metrics, episode cost sums and the near-target fraction on a deterministic test system, the appended identification data, the local Jacobian on the linear system, the LQR reference, the action-cost scale and both guards. 14 tests, about 4 seconds.koopmanrl_utils/AGENTS.mdlists both modules;.gitignoreignores both output directories.Code comments
The comments describe each computation in full, so the code can be read without the paper: the normalised one-step error is written out, the R·dt action-cost scaling is explained, and the comments say why appended transitions are rounded to single precision and why the gain uses the generator state left by the previous evaluation. Where the code implements a theoretical statement, the docstring gives the formula and cites the statement by number and name:
action_fields(Proposition S11.1, the policy is an exponential family over the action dictionary) andpolicy_reading(Corollary S11.2, the Gaussian closed form). Each module docstring ends with a single reference to the paper's supplementary material, by section title. Tables, figures and section numbers are no longer named in the code, and the summary prints "Accuracy" and "Sensitivity" in place of table numbers. Apart from those two printed labels, 1cebcf5 changes no code: with docstrings removed, the syntax trees of both modules are otherwise identical.Running it
No new dependencies.
Verification. The S16 tables were first produced by exploratory scripts that
skvi_sensitivity_checks.pyconsolidates. A full run of the module on all eight seeds agrees exactly with the saved results of those scripts for every reported number of both studies (1,624 values), and--summarize_onlyprints the medians of Tables S6 to S8. The re-identification ablation, run on the same eight seeds, gives a median Lorenz cost of 97,000 with the tensor refitted only and 844,000 with the policy states added to the value fit only, against 241,000 for the tuned SKVI and 48,000 for full re-identification. The exception is the model Jacobians in the sensitivity study, which the module evaluates in double precision with a smaller step (differences below 1e-7). An independent review by Codex (GPT, extra-high reasoning) found no error in the metrics, the conditional means, the paired noise, the LQR reference or the cost weighting. Its seven requests on clarity, provenance and tests are addressed in the S16 commits. Ruff, ruff format, isort, vulture and the hygiene hooks pass.🤖 Generated with Claude Code