Skip to content

Fix Codex limits collector for CLI approval-policy churn - #8977

Closed
fresh3nough wants to merge 6 commits into
omacom:quattrofrom
fresh3nough:fix/issue-8971-codex-approval-policy
Closed

fresh3nough wants to merge 6 commits into
omacom:quattrofrom
fresh3nough:fix/issue-8971-codex-approval-policy

Conversation

@fresh3nough

Copy link
Copy Markdown
Contributor

Problem

With codex-cli >= 0.149 (reporter: 0.151.0), the agents panel shows CODEX LIMITS UNAVAILABLE and the auth-help card reads the bare word initialize. Token stats still render; only limits/plan are missing.

bin/omarchy-agent-usage-codex used to spawn:

codex -s read-only -a untrusted app-server

Current CLI rejects untrusted (possible values: on-request, never), exits immediately, and rpc_request() raises TimeoutError("initialize"), which was stored verbatim as authHelpText.

Related: #8849 #8868 #8721 #8656. Primary flag change already landed on quattro as #7649 (on-request); this PR finishes the residual failure mode (early exit still looked like a hang / leaked method name) and prefers never for a non-interactive probe that never starts a turn.

Fix

  • Spawn with -a never (valid on current CLI; same "do not prompt" intent as the old untrusted).
  • Detect app-server exit during send/read immediately (rpc_send / rpc_request) instead of waiting out the timeout.
  • Capture stderr on failure and prefer the CLI's own message over bare method names; fall back to the login hint only when the process exits silently.
  • Timeouts on a live process become Codex app-server did not answer <method>.

Verification

  • bash test/shell.d/agent-usage-codex-scanner-test.sh (26 ok)
  • bash test/shell.d/agent-usage-update-test.sh
  • python3 -m py_compile bin/omarchy-agent-usage-codex
  • GCE Debian 12 worker, PATH-isolated collector:
collector result
v4.0.1 (-a untrusted) usageStatusText: Codex limits unavailable, authHelpText: initialize, limits: []
this branch (-a never + exit detection) tierLabel: pro, weekly limit 8%, empty status

Notes

Stable tag v4.0.1 still has untrusted. v4-0-2 / quattro already had on-request from #7649; this is additive resilience + clearer panel text.

Fixes #8971

codex-cli >= 0.149 dropped -a untrusted. Launch the app-server with
-a never (non-interactive, no prompts), detect early process exit
instead of timing out on initialize, and surface CLI stderr rather
than the bare RPC method name in the agents panel.

Fixes omacom#8971

Signed-off-by: anonwurcod <anonwurcod@proton.me>
omarchybot and others added 2 commits October 1, 2026 16:49
quattro now asks account/rateLimits/read first and makes account/read an optional fallback that swallows a timeout. An app-server that exits during that fallback now raises RuntimeError rather than TimeoutError, so the fallback catches both, or the limits already read would be lost. The stall test moves to account/rateLimits/read, the request quattro still waits on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
codex_rpc_help read stderr before checking whether the process was still running, so a stalled app-server that had written any warning was reported as "codex app-server exited: <the warning>". Stderr only explains a failure once the process has exited; while it is up, the stall is the finding.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@omarchybot omarchybot added the verified Omarchy Triage has verified that this issue is ready for final review label Oct 1, 2026
@omarchybot

Copy link
Copy Markdown
Collaborator

Reviewed and verified at e21b1c0 against quattro (c05d901), on a disposable Omarchy worker.

Reproduction. I used a stand-in codex that rejects its arguments the way codex-cli 0.151 rejected -a untrusted. With quattro's collector the panel record is authHelpText: "initialize". With this branch it is the CLI's own message: codex app-server exited: error: invalid value 'untrusted' for '--ask-for-approval <APPROVAL_POLICY>' [possible values: on-request, never]. With the collector reverted, the new scanner test fails on exactly that initialize.

Pushed to the branch. I pushed two commits:

  • 4e2d7ce merges quattro in, because the branch no longer merged. quattro now asks account/rateLimits/read first and keeps account/read only as an optional fallback that ignores a TimeoutError. This branch raises RuntimeError when the app-server exits, so the fallback now catches both. Otherwise an app-server exiting during the fallback would throw away the limits already read. The stall test now stalls on account/rateLimits/read, the request quattro still waits on.
  • e21b1c0 fixes codex_rpc_help, which read stderr before checking whether the process was still running. A stalled app-server that had logged anything, a warning for instance, was reported as codex app-server exited: <the warning>. It now reports a live process as a stall and uses stderr only after an exit. A new test covers this; it fails without the change.

Tests on the worker. agent-usage-codex-scanner-test.sh passed 30 of 30, agent-usage-update-test.sh 7 of 7, agent-usage-accounts-test.sh 11 of 11 and agents-panel-test.sh 24 of 24. ./test/cli stops at vscode generated theme references current theme file, which fails the same way on quattro and has nothing to do with this change. That means the rest of the CLI suite did not run on either tree. The PR has no CI checks.

Scope. Switching -a on-request to -a never is not needed for this fix, since quattro's on-request is accepted by current codex-cli. It is a one-line choice, so I left it for the maintainer.

Related. #13740 fixes the same symptom by replacing the help text with generic messages. This PR does more: it reports the CLI's stderr and tells an app-server that exited apart from one that stalled. Under #13740, a CLI that rejects its arguments would read "didn't answer in time, will retry", which is wrong. #13703 and #13816 change the same rpc_request reader to fix a different bug, replies stuck in Python's read buffer, so whichever lands second will conflict with this one. #8971, which this PR closes, was already closed as a duplicate of #8460, and quattro already fixes the stale flag that #8460 reports.

Who checked it. Claude Opus 5.5 only. The second opinion (Codex Medium) did not run: it answered a ping, but the review was refused because its daily budget was used up. Nothing here has had a second review.

This waits on the maintainer: a second review, then a choice between this and #13740.

@greptile-apps

greptile-apps Bot commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Improves error handling in the Codex CLI integration.

The PR appears safe to merge, with non-blocking diagnostic and test-reliability issues remaining.

Findings

  1. P2 Startup errors become login hints ▶
  2. P2 Test depends on timing ▶

Summary

The PR switches the non-interactive Codex probe to -a never and improves app-server exit diagnostics. Since the previous review, it also adds a temporary-file failure fallback, retains the end of stderr, and extends two test fixtures’ lifetimes.

Reviews (2) · Last reviewed commit: "Keep the exit-path fixtures alive until ..."

Comment thread bin/omarchy-agent-usage-codex Outdated
Comment thread bin/omarchy-agent-usage-codex Outdated
Comment on lines +675 to +676
exec 1>&-
sleep 2

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Test depends on timing This fixture closes stdout but stays alive for only two seconds, while the assertion expects the collector to inspect it before it exits. If a runner is delayed, the collector returns the silent-exit login hint instead and the test fails despite unchanged behavior. Make the fixture's lifetime independent of runner timing.

omarchybot and others added 3 commits October 1, 2026 17:14
The stderr temp file was created outside the guarded launch, so an unwritable temp directory or exhausted descriptors raised out of fetch_codex_rpc and the collector printed no record at all, leaving the panel on its last one. Fall back to discarding stderr instead: the probe still runs, and only the CLI's own error text is lost.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The exit message joined stderr from the top and cut it at 300 characters, so a CLI that logged a few warnings on the way up had its actual error cut off and the panel showed the warnings instead. Keep the tail.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both fixtures stayed up for only two seconds, while the collector gives a failed process up to a second to exit before deciding whether it is still running. A slow runner could see them exit first and get the login hint instead of the message under test. They now live until the collector terminates them, which it always does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@omarchybot omarchybot removed the verified Omarchy Triage has verified that this issue is ready for final review label Oct 1, 2026
@omarchybot

Copy link
Copy Markdown
Collaborator

Followed up on Greptile's three comments at e21b1c0 and pushed a fix for each. All three held up against the source.

  • 61c1c72 handles the case where the temp file for the CLI's stderr cannot be created, such as an unwritable temp directory or no file descriptors left. That file was created outside the guarded launch, so the error escaped fetch_codex_rpc and the collector printed no record at all, leaving the panel on its old one. It now discards stderr instead and still probes. On a worker, with tempfile.TemporaryFile forced to raise, e21b1c0 exits with a traceback and this head prints a full record.
  • 72431cd keeps the end of stderr rather than the start, because the fatal error is the last thing the CLI writes. With twenty warning lines ahead of the error, e21b1c0's panel text was all warnings; this head's ends in the error. A new test covers this: it fails against e21b1c0's collector and passes here.
  • 6b980ae fixes two exit-path fixtures that stayed up for only two seconds. The collector gives a failed process up to a second before checking whether it is still running, so on a slow runner they could exit first and the tests would get the login hint instead. They now run until the collector terminates them.

Tested on a disposable Omarchy worker at 6b980ae against quattro (c05d901). agent-usage-codex-scanner-test.sh passed 31 of 31, agent-usage-update-test.sh 7 of 7, agent-usage-accounts-test.sh 11 of 11 and agents-panel-test.sh 24 of 24. ./test/cli still stops at vscode generated theme references current theme file, which fails the same way on quattro, so the rest of the CLI suite did not run on either tree.

Checked by Claude Opus 5.5 alone. The second opinion (Codex Medium) answered a ping, but the review was refused again because today's review budget is used up, so this change has still not had a second review.

This still waits on the maintainer for the same things as before: a second review, the -a never switch, and a choice between this and #13740. #13703 and #13816 still touch the same rpc_request reader.

stdin=subprocess.PIPE,
stdout=subprocess.PIPE,
stderr=subprocess.DEVNULL,
stderr=stderr_file or subprocess.DEVNULL,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Startup errors become login hints If the temporary file cannot be created, this fallback discards Codex’s error output. If Codex then exits with a startup error, the collector tells the user to run codex login instead of showing the cause, making unavailable limits harder to diagnose.

@omarchybot omarchybot added the verified Omarchy Triage has verified that this issue is ready for final review label Oct 1, 2026
@omarchybot

Copy link
Copy Markdown
Collaborator

Greptile's newest comment, that a startup error turns into the codex login hint when no temp file can be made, is accurate. In that case stderr_file is None, codex_rpc_help finds no stderr to show, and it falls back to AUTH_HELP. It only happens when two things fail together: the temp directory is unwritable or out of file descriptors, and Codex then exits on startup. The commit before it (61c1c72) made that case print a record at all instead of crashing, so I have not pushed another change to the branch for it. If the maintainer wants it, the smallest change is to have codex_rpc_help report the exit status rather than the login hint when stderr_file is None.

6b980ae is unchanged since my last comment, so the worker results there still stand, and this head is labelled verified again.

Checked by Claude Opus 5.5 alone. The second opinion (Codex Medium) answered a ping, but the review was refused again because today's 500 reviews are used up, so nothing on this branch has had a second review.

This still waits on the maintainer: a second review, the -a never switch, and a choice between this and #13740.

@dhh

dhh commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Thank you! Brought in with your authorship in #14049.

@dhh dhh closed this Oct 2, 2026
dhh added a commit that referenced this pull request Oct 3, 2026
…ommunity PRs (#14049)

* Read Codex app-server replies from the raw fd (#13703)

* Resolve Codex through mise which instead of running the lazy launcher (#13109)

* Skip the Codex app-server probe when there are no credentials (#13106)

Adapted: credentials are checked in the home being probed rather than in
the CODEX_HOME environment variable, since each registered account is
probed in its own home, so a signed-out secondary account isn't hidden
behind the primary's login. A home without credentials reports "Waiting
for auth" like any other signed-out home. The credentials store setting is
read with tomllib, so a single-quoted value counts too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Show the Codex CLI's own error when its app-server dies (#8977)

Detect an app-server that exits or stops answering, and report the end of
its stderr instead of a bare RPC method name. Rebased onto the raw-fd
reply reader; the switch from "-a on-request" to "-a never" is left out,
keeping the current approval flags.

* Count pi sessions when HOME is a git checkout (#13209)

* Count only OpenAI-backed native sessions as Codex usage (#12032)

* Deduplicate Pi usage across forked sessions (#8602)

* Skip unchanged native Codex token snapshots (#10531)

* Count omp and pi profile sessions in the agent usage collectors (#9546)

`omp --profile=<name>` (and pi's equivalent) relocates the whole agent
tree under <base>/profiles/<name>/. The Claude and Codex collectors only
ever scanned <base>/agent/sessions, so a subscription driven entirely
through a profile was invisible to the agents panel: no tokens by day, no
tokens by model, no prompt or session counts.

Discover the profile roots alongside the default one. Sessions are keyed
by file path, so a profile adds sessions instead of double-counting the
default root, and a missing or unreadable profiles directory leaves the
existing behavior untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014zFbJcDEEpV5BAmsH6kAB3

* Skip unrelated Codex session lines before JSON parsing (#12803)

Adapted: session_meta lines also pass the pre-filter, since the provider
filter from #12032 reads them to skip rollouts served by a non-OpenAI
provider.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read only the Codex session files that changed since the last scan (#12595)

Native Codex rollouts keep per-file totals between runs, replayed while a
file's mtime and size are unchanged. Rebased onto the session_meta
provider filter, snapshot dedup, and line pre-filter, which now live in
the per-file reader. pi and omp sessions are left out of the per-file
cache: a forked pi session repeats its parent's messages, so they are
deduplicated across the whole tree on every scan.

* Count streamed Claude messages by their highest-output usage line (#10606)

Claude Code writes a streamed assistant response as several transcript
lines that share one message id, one per content block. Each line
carries a usage object. The first line's output_tokens is a placeholder,
often 1, and the last line has the real count. Input and cache fields
usually match across the lines.

The scanner dedupes by message id and keeps the first line it sees, so
it under-counts output tokens. On a machine with 2,577 transcripts it
reported 39.0M output tokens against 60.1M used, a 35% shortfall. Input
and both cache fields differed by under 0.01%.

Keep the line with the highest output count, with the last one scanned
winning a tie. The whole line is kept because a response can fall back
to another model mid-stream. Those lines are separate snapshots with
different cache figures and a different model, and taking a maximum per
field across them over-counts cache tokens and credits the wrong model.

The zero-usage check now runs before dedup, so a zero-usage first line
no longer claims a message id and hides a later line with real usage.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: GPT-6 Astra <noreply@openai.com>

* Index Claude transcripts so the agents refresh reads only what was appended (#8313)

omarchy-agent-usage-claude re-parsed every line of every transcript under
~/.claude/projects on each refresh: no mtime cutoff, no memory of the last
pass. The agents widget is on by default and ticks every 15 minutes, so the
cost grew for the life of the machine. After one month here that was 803
files, 640 MB, 127k lines and 57k JSON parses per tick, about 1 core-second,
pushed through the page cache every quarter hour forever.

Keep a per-file index next to the scan cache: the unique usage records
already parsed out of each transcript and the byte offset they end at. A
file whose size and mtime match is not opened; a file that grew is read
from the stored offset; a file that shrank or was rewritten is read from
the start. --force drops the index and rescans from scratch.

The summary is built from the indexed records in the same directory order
the walk always used. That matters: when a resumed session carries earlier
messages, the same message id appears in two files with different usage,
and the first file visited wins. 91 ids differed on this machine; sorting
the walk moved one model's output total by 25k tokens. Output is now
byte-identical to the previous scan on a frozen copy of the corpus, cold,
warm, and after an append.

Warm refresh: 1.0 s -> 0.10 s of CPU, of which the scan itself is 70 ms;
the index for this corpus is 2.9 MB.

Adapted:
- Rebased onto #10606: the highest-output rule for streamed messages now
  lives where the index parses records, and decides between files too.
- The index records the timezone it was written in, and a change rereads
  every transcript, since its records hold local days.
- A file only counts as appended to when its inode and the hash of what
  was already read still match, so a transcript replaced by a larger one,
  or rewritten in place, is read from the start.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Label a Claude Team seat by its subscription, not its rate-limit tier (#11109)

The collector built the plan label from the OAuth rateLimitTier first, so a
Team premium seat, which runs on default_claude_max_5x, showed in the agents
panel as "Max 5x". Lead with subscriptionType and keep the multiplier as its
qualifier: Max still reads "Max 5x", a Team seat reads "Team 5x".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Label the Claude plan from the profile the CLI refreshes (#7225)

Adapted: the profile is found the same way current_account_id() finds it,
now shared as profile_path(): ~/.claude.json for the default home, the
home's own .claude.json otherwise. The original fell back to ~/.claude.json
for any home without CLAUDE_CONFIG_DIR set, so a secondary account read the
primary's tier. The profile's tier also keeps the subscription in the label,
so a Team seat stays "Team" (#11109).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Call a lapsed Claude access token paused, not signed out (#8093)

* Refresh Claude usage after the clock moves backwards (#9956)

* Bound unreadable Claude transcript warnings (#12414)

* Count Claude usage from opencode v2 sessions (#13894)

* Reload agent usage records when an inotify watch fails to rearm (#10067)

* Reload agent usage records after each update run instead of on a timer

Rather than #10067's two-minute timer per record, reload every record when
the omarchy-agent-usage-update process exits, the moment its files can have
been replaced. A reload that finds a file unchanged keeps its record, so the
panel isn't stirred up by identical data. The grep test now runs the QML
functions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Show the agent status when the trouble line has no help text (#8497)

* Clear stale agent login guidance after a successful probe (#8892)

* Clear the Grok login hint after a successful probe

#8892 cleared the default login hint after a successful probe in the
Claude and Codex collectors; Grok's collector had the same stale hint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read Fireworks credentials from pi's auth.json (#7455)

The Fireworks collector skipped pi, Omarchy's default agent, when
walking its credential ladder, so a machine signed in to Fireworks only
through pi (/login fireworks) never showed the tab. Insert the key pi
stores in $PI_CODING_AGENT_DIR/auth.json (default ~/.pi/agent) between
the firectl auth.ini and the opencode fallback.

pi keys can be literals, $ENV_VAR/${ENV_VAR} references, or !command
shell lookups. The collector resolves the first two; command lookups
stay pi-only and are skipped rather than sent to the API verbatim.

* Call a lapsed Grok access token paused, not signed out

Grok's access token lives six hours and Grok mints a new one from its
refresh token whenever it starts, so a lapsed one is routine. Reporting
it as an expired sign-in made the panel offer Sign-in required several
times a day, sending people through grok login for nothing. With a
refresh token present it now reads as paused, like Claude's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep showing Grok's last limits while it sits idle

While Grok hasn't run, nothing on the machine has spent its allowance,
so with a refresh token on hand the last numbers still stand: they show
as current rather than dimmed under a status line. A weekly window that
reset in the meantime starts over at 0%, a whole number of weeks on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Ask for a Grok sign-in once its refresh token is past 30 days

A refresh token older than Grok's 30-day sign-in can't renew anything,
so the panel offers Sign-in required again instead of showing the last
limits as current. With nothing cached yet it says to start Grok, rather
than showing an empty section without a word.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Check both ends of what the Claude index read before resuming a transcript

A transcript rewritten in place could grow and change only after its
first kilobytes, and the index took it for an append. It now compares
the last kilobytes before the resume point too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Simplify the agent usage collectors

- Codex: pass the forced-scan choice down instead of a module global, make
  the per-file reader's cache arguments required, shrink the cache record
  check, and drop guards for shapes that can't occur: an empty launcher
  path, realpath raising, mise itself being a lazy launcher, multi-line
  `mise which` output, and probing without a temp file for stderr.
- Claude: decide an append by the digest of both ends of what was read
  alone; the inode and mtime checks it made redundant are gone.
- Snapshot: the device id falls back to the hostname, which always exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Share fixture setup in the agent usage scanner tests

Every fixture home lives under one scratch directory with a single cleanup
trap, instead of a trap rewritten with a longer list for each new home, and
the Codex test builds its signed-in homes with one helper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Treat a replaced Claude transcript as new even when its ends match

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Probe Codex without its error text when there's no temporary space

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: tossbaws <17258053+tossbaws@users.noreply.github.com>
Co-authored-by: surim0n <suritech@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: anonwurcod <anonwurcod@proton.me>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Co-authored-by: Nate Ashby <nate.ashby11@gmail.com>
Co-authored-by: Aris Gysel <aris.gysel@me.com>
Co-authored-by: Brams <76213579+Brams-s@users.noreply.github.com>
Co-authored-by: This_Is_NPC <gabrielfollone27@gmail.com>
Co-authored-by: sanjyay <102979855+sanjyay@users.noreply.github.com>
Co-authored-by: PapistProtocol <12738904+PapistProtocol@users.noreply.github.com>
Co-authored-by: steez <stevedimakos97@gmail.com>
Co-authored-by: GPT-6 Astra <noreply@openai.com>
Co-authored-by: Ryan Yogan <ryanyogan@gmail.com>
Co-authored-by: Oli Denton <41393837+omdenton@users.noreply.github.com>
Co-authored-by: Igor Kramar <i@ikramar.ru>
Co-authored-by: Martin Eidensten <martin@meibe.se>
Co-authored-by: Romain Perron <rdj.perron@gmail.com>
Co-authored-by: Omarchy Contributor <contributor@users.noreply.github.com>
Co-authored-by: manuaudio <manu@arimaka.com>
Co-authored-by: Tyler South <tsouth2@gmail.com>
Co-authored-by: whathek <Hek846@users.noreply.github.com>
Co-authored-by: Ty Richards <me@tyrichards.com>
sgruendel pushed a commit to sgruendel/omarchy that referenced this pull request Oct 3, 2026
…ommunity PRs (omacom#14049)

* Read Codex app-server replies from the raw fd (omacom#13703)

* Resolve Codex through mise which instead of running the lazy launcher (omacom#13109)

* Skip the Codex app-server probe when there are no credentials (omacom#13106)

Adapted: credentials are checked in the home being probed rather than in
the CODEX_HOME environment variable, since each registered account is
probed in its own home, so a signed-out secondary account isn't hidden
behind the primary's login. A home without credentials reports "Waiting
for auth" like any other signed-out home. The credentials store setting is
read with tomllib, so a single-quoted value counts too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Show the Codex CLI's own error when its app-server dies (omacom#8977)

Detect an app-server that exits or stops answering, and report the end of
its stderr instead of a bare RPC method name. Rebased onto the raw-fd
reply reader; the switch from "-a on-request" to "-a never" is left out,
keeping the current approval flags.

* Count pi sessions when HOME is a git checkout (omacom#13209)

* Count only OpenAI-backed native sessions as Codex usage (omacom#12032)

* Deduplicate Pi usage across forked sessions (omacom#8602)

* Skip unchanged native Codex token snapshots (omacom#10531)

* Count omp and pi profile sessions in the agent usage collectors (omacom#9546)

`omp --profile=<name>` (and pi's equivalent) relocates the whole agent
tree under <base>/profiles/<name>/. The Claude and Codex collectors only
ever scanned <base>/agent/sessions, so a subscription driven entirely
through a profile was invisible to the agents panel: no tokens by day, no
tokens by model, no prompt or session counts.

Discover the profile roots alongside the default one. Sessions are keyed
by file path, so a profile adds sessions instead of double-counting the
default root, and a missing or unreadable profiles directory leaves the
existing behavior untouched.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014zFbJcDEEpV5BAmsH6kAB3

* Skip unrelated Codex session lines before JSON parsing (omacom#12803)

Adapted: session_meta lines also pass the pre-filter, since the provider
filter from omacom#12032 reads them to skip rollouts served by a non-OpenAI
provider.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read only the Codex session files that changed since the last scan (omacom#12595)

Native Codex rollouts keep per-file totals between runs, replayed while a
file's mtime and size are unchanged. Rebased onto the session_meta
provider filter, snapshot dedup, and line pre-filter, which now live in
the per-file reader. pi and omp sessions are left out of the per-file
cache: a forked pi session repeats its parent's messages, so they are
deduplicated across the whole tree on every scan.

* Count streamed Claude messages by their highest-output usage line (omacom#10606)

Claude Code writes a streamed assistant response as several transcript
lines that share one message id, one per content block. Each line
carries a usage object. The first line's output_tokens is a placeholder,
often 1, and the last line has the real count. Input and cache fields
usually match across the lines.

The scanner dedupes by message id and keeps the first line it sees, so
it under-counts output tokens. On a machine with 2,577 transcripts it
reported 39.0M output tokens against 60.1M used, a 35% shortfall. Input
and both cache fields differed by under 0.01%.

Keep the line with the highest output count, with the last one scanned
winning a tie. The whole line is kept because a response can fall back
to another model mid-stream. Those lines are separate snapshots with
different cache figures and a different model, and taking a maximum per
field across them over-counts cache tokens and credits the wrong model.

The zero-usage check now runs before dedup, so a zero-usage first line
no longer claims a message id and hides a later line with real usage.

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
Co-authored-by: GPT-6 Astra <noreply@openai.com>

* Index Claude transcripts so the agents refresh reads only what was appended (omacom#8313)

omarchy-agent-usage-claude re-parsed every line of every transcript under
~/.claude/projects on each refresh: no mtime cutoff, no memory of the last
pass. The agents widget is on by default and ticks every 15 minutes, so the
cost grew for the life of the machine. After one month here that was 803
files, 640 MB, 127k lines and 57k JSON parses per tick, about 1 core-second,
pushed through the page cache every quarter hour forever.

Keep a per-file index next to the scan cache: the unique usage records
already parsed out of each transcript and the byte offset they end at. A
file whose size and mtime match is not opened; a file that grew is read
from the stored offset; a file that shrank or was rewritten is read from
the start. --force drops the index and rescans from scratch.

The summary is built from the indexed records in the same directory order
the walk always used. That matters: when a resumed session carries earlier
messages, the same message id appears in two files with different usage,
and the first file visited wins. 91 ids differed on this machine; sorting
the walk moved one model's output total by 25k tokens. Output is now
byte-identical to the previous scan on a frozen copy of the corpus, cold,
warm, and after an append.

Warm refresh: 1.0 s -> 0.10 s of CPU, of which the scan itself is 70 ms;
the index for this corpus is 2.9 MB.

Adapted:
- Rebased onto omacom#10606: the highest-output rule for streamed messages now
  lives where the index parses records, and decides between files too.
- The index records the timezone it was written in, and a change rereads
  every transcript, since its records hold local days.
- A file only counts as appended to when its inode and the hash of what
  was already read still match, so a transcript replaced by a larger one,
  or rewritten in place, is read from the start.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Label a Claude Team seat by its subscription, not its rate-limit tier (omacom#11109)

The collector built the plan label from the OAuth rateLimitTier first, so a
Team premium seat, which runs on default_claude_max_5x, showed in the agents
panel as "Max 5x". Lead with subscriptionType and keep the multiplier as its
qualifier: Max still reads "Max 5x", a Team seat reads "Team 5x".

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* Label the Claude plan from the profile the CLI refreshes (omacom#7225)

Adapted: the profile is found the same way current_account_id() finds it,
now shared as profile_path(): ~/.claude.json for the default home, the
home's own .claude.json otherwise. The original fell back to ~/.claude.json
for any home without CLAUDE_CONFIG_DIR set, so a secondary account read the
primary's tier. The profile's tier also keeps the subscription in the label,
so a Team seat stays "Team" (omacom#11109).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Call a lapsed Claude access token paused, not signed out (omacom#8093)

* Refresh Claude usage after the clock moves backwards (omacom#9956)

* Bound unreadable Claude transcript warnings (omacom#12414)

* Count Claude usage from opencode v2 sessions (omacom#13894)

* Reload agent usage records when an inotify watch fails to rearm (omacom#10067)

* Reload agent usage records after each update run instead of on a timer

Rather than omacom#10067's two-minute timer per record, reload every record when
the omarchy-agent-usage-update process exits, the moment its files can have
been replaced. A reload that finds a file unchanged keeps its record, so the
panel isn't stirred up by identical data. The grep test now runs the QML
functions.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Show the agent status when the trouble line has no help text (omacom#8497)

* Clear stale agent login guidance after a successful probe (omacom#8892)

* Clear the Grok login hint after a successful probe

omacom#8892 cleared the default login hint after a successful probe in the
Claude and Codex collectors; Grok's collector had the same stale hint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Read Fireworks credentials from pi's auth.json (omacom#7455)

The Fireworks collector skipped pi, Omarchy's default agent, when
walking its credential ladder, so a machine signed in to Fireworks only
through pi (/login fireworks) never showed the tab. Insert the key pi
stores in $PI_CODING_AGENT_DIR/auth.json (default ~/.pi/agent) between
the firectl auth.ini and the opencode fallback.

pi keys can be literals, $ENV_VAR/${ENV_VAR} references, or !command
shell lookups. The collector resolves the first two; command lookups
stay pi-only and are skipped rather than sent to the API verbatim.

* Call a lapsed Grok access token paused, not signed out

Grok's access token lives six hours and Grok mints a new one from its
refresh token whenever it starts, so a lapsed one is routine. Reporting
it as an expired sign-in made the panel offer Sign-in required several
times a day, sending people through grok login for nothing. With a
refresh token present it now reads as paused, like Claude's.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep showing Grok's last limits while it sits idle

While Grok hasn't run, nothing on the machine has spent its allowance,
so with a refresh token on hand the last numbers still stand: they show
as current rather than dimmed under a status line. A weekly window that
reset in the meantime starts over at 0%, a whole number of weeks on.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Ask for a Grok sign-in once its refresh token is past 30 days

A refresh token older than Grok's 30-day sign-in can't renew anything,
so the panel offers Sign-in required again instead of showing the last
limits as current. With nothing cached yet it says to start Grok, rather
than showing an empty section without a word.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Check both ends of what the Claude index read before resuming a transcript

A transcript rewritten in place could grow and change only after its
first kilobytes, and the index took it for an append. It now compares
the last kilobytes before the resume point too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Simplify the agent usage collectors

- Codex: pass the forced-scan choice down instead of a module global, make
  the per-file reader's cache arguments required, shrink the cache record
  check, and drop guards for shapes that can't occur: an empty launcher
  path, realpath raising, mise itself being a lazy launcher, multi-line
  `mise which` output, and probing without a temp file for stderr.
- Claude: decide an append by the digest of both ends of what was read
  alone; the inode and mtime checks it made redundant are gone.
- Snapshot: the device id falls back to the hostname, which always exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Share fixture setup in the agent usage scanner tests

Every fixture home lives under one scratch directory with a single cleanup
trap, instead of a trap rewritten with a longer list for each new home, and
the Codex test builds its signed-in homes with one helper.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Treat a replaced Claude transcript as new even when its ends match

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Probe Codex without its error text when there's no temporary space

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: tossbaws <17258053+tossbaws@users.noreply.github.com>
Co-authored-by: surim0n <suritech@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: anonwurcod <anonwurcod@proton.me>
Co-authored-by: Kevin Rajan <7121943+kvnloo@users.noreply.github.com>
Co-authored-by: Nate Ashby <nate.ashby11@gmail.com>
Co-authored-by: Aris Gysel <aris.gysel@me.com>
Co-authored-by: Brams <76213579+Brams-s@users.noreply.github.com>
Co-authored-by: This_Is_NPC <gabrielfollone27@gmail.com>
Co-authored-by: sanjyay <102979855+sanjyay@users.noreply.github.com>
Co-authored-by: PapistProtocol <12738904+PapistProtocol@users.noreply.github.com>
Co-authored-by: steez <stevedimakos97@gmail.com>
Co-authored-by: GPT-6 Astra <noreply@openai.com>
Co-authored-by: Ryan Yogan <ryanyogan@gmail.com>
Co-authored-by: Oli Denton <41393837+omdenton@users.noreply.github.com>
Co-authored-by: Igor Kramar <i@ikramar.ru>
Co-authored-by: Martin Eidensten <martin@meibe.se>
Co-authored-by: Romain Perron <rdj.perron@gmail.com>
Co-authored-by: Omarchy Contributor <contributor@users.noreply.github.com>
Co-authored-by: manuaudio <manu@arimaka.com>
Co-authored-by: Tyler South <tsouth2@gmail.com>
Co-authored-by: whathek <Hek846@users.noreply.github.com>
Co-authored-by: Ty Richards <me@tyrichards.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working verified Omarchy Triage has verified that this issue is ready for final review

Projects

None yet

Development

Successfully merging this pull request may close these issues.

agents: Codex limits break with codex-cli >= 0.151.0 (invalid '-a untrusted')

3 participants