Repository navigation
feat(providers): add Nemotron 3 Ultra, Nemotron 3.5 Lightning and GLM 5.3 via OpenRouter to the model matrix - #716
Merged
Conversation
…odel matrix (OPE-215)
…st (OPE-215) The window is the Z.AI endpoint's own 1,048,576 (fp8, 131,072 max output), read from openrouter.ai/api/v1/models/z-ai/glm-5.3/endpoints on 2026-10-01, because Z.AI is the host a pinned run uses. Without a row the model falls back to DEFAULT_CONTEXT_WINDOW (128,000) and compacts at 102,400, far inside what the endpoint serves, so any comparison against a model with a correct row would partly measure compaction. The matrix size guard moves from 70 to 75 for the three OpenRouter rows added under OPE-215 (Nemotron 3 Ultra, Nemotron 3.5 Lightning, GLM 5.3). The pruning the guard's note asks for before another raise is still owed and stays an owner call.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What and why. Three OpenRouter rows for models that had no entry, so they ran with the 128k fallback window and compacted at 102k regardless of what their endpoint serves.
The windows. Each row carries the figure from the host a pinned run uses, with the date it was read, so compaction triggers inside the real window where hosts disagree:
openrouter:nvidia/nemotron-3-ultra-550b-a55b— 202,800 (BaseTen)openrouter:nvidia/nemotron-3.5-lightning— 262,144 (all endpoints agree)openrouter:z-ai/glm-5.3— 1,048,576 (Z.AI)Read from
openrouter.ai/api/v1/models/<id>/endpointson 2026-10-01. Paid endpoints only; the:freevariants are left out because they log and may train on inputs.Tests. The same three rows are added to
tests/test_matrix_current_models.pyso the figures are checked.Guard raise, 70 → 75. The rows take the matrix to 70. The guard's note asks for retired rows to be pruned before the next raise. Pruning is kept out of this PR on purpose: dropping a row silently downgrades that model to the fallback capabilities, so it deserves its own change with its own review rather than riding along with three additions. The note stays, with a line recording this raise.
Not in this PR. Price rows live in
coworker/headless/prices.yaml, which feat(cli): addopenworker runfor unattended tasks, adangerously-bypass-approvalsmode and anautoattendance (OPE-196, OPE-203, OPE-204) #680 introduces; they follow after it merges.Commits and tests. Two commits;
tests/test_providers.pyandtests/test_matrix_current_models.py, 57 passed.