Skip to content

feat(providers): add Nemotron 3 Ultra, Nemotron 3.5 Lightning and GLM 5.3 via OpenRouter to the model matrix - #716

Merged
rohitprasad15 merged 2 commits into
mainfrom
issue/ope-215-nemotron-matrix
Oct 5, 2026
Merged

rohitprasad15 merged 2 commits into
mainfrom
issue/ope-215-nemotron-matrix

Conversation

@devikaverma

Copy link
Copy Markdown
Collaborator
  • What and why. Three OpenRouter rows for models that had no entry, so they ran with the 128k fallback window and compacted at 102k regardless of what their endpoint serves.

  • The windows. Each row carries the figure from the host a pinned run uses, with the date it was read, so compaction triggers inside the real window where hosts disagree:

    • openrouter:nvidia/nemotron-3-ultra-550b-a55b — 202,800 (BaseTen)
    • openrouter:nvidia/nemotron-3.5-lightning — 262,144 (all endpoints agree)
    • openrouter:z-ai/glm-5.3 — 1,048,576 (Z.AI)

    Read from openrouter.ai/api/v1/models/<id>/endpoints on 2026-10-01. Paid endpoints only; the :free variants are left out because they log and may train on inputs.

  • Tests. The same three rows are added to tests/test_matrix_current_models.py so the figures are checked.

  • Guard raise, 70 → 75. The rows take the matrix to 70. The guard's note asks for retired rows to be pruned before the next raise. Pruning is kept out of this PR on purpose: dropping a row silently downgrades that model to the fallback capabilities, so it deserves its own change with its own review rather than riding along with three additions. The note stays, with a line recording this raise.

  • Not in this PR. Price rows live in coworker/headless/prices.yaml, which feat(cli): add openworker run for unattended tasks, a dangerously-bypass-approvals mode and an auto attendance (OPE-196, OPE-203, OPE-204) #680 introduces; they follow after it merges.

  • Commits and tests. Two commits; tests/test_providers.py and tests/test_matrix_current_models.py, 57 passed.

…st (OPE-215)

The window is the Z.AI endpoint's own 1,048,576 (fp8, 131,072 max output), read from
openrouter.ai/api/v1/models/z-ai/glm-5.3/endpoints on 2026-10-01, because Z.AI is the
host a pinned run uses. Without a row the model falls back to DEFAULT_CONTEXT_WINDOW
(128,000) and compacts at 102,400, far inside what the endpoint serves, so any
comparison against a model with a correct row would partly measure compaction.

The matrix size guard moves from 70 to 75 for the three OpenRouter rows added under
OPE-215 (Nemotron 3 Ultra, Nemotron 3.5 Lightning, GLM 5.3). The pruning the guard's
note asks for before another raise is still owed and stays an owner call.

@rohitprasad15 rohitprasad15 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@rohitprasad15
rohitprasad15 merged commit 5d6c9e8 into main Oct 5, 2026
6 checks passed
@rohitprasad15
rohitprasad15 deleted the issue/ope-215-nemotron-matrix branch October 5, 2026 21:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants