Skip to content

Fix Codex usage timeouts on batched replies - #12979

Closed
orienw wants to merge 1 commit into
omacom:quattrofrom
orienw:fix/codex-usage-stdout
Closed

orienw wants to merge 1 commit into
omacom:quattrofrom
orienw:fix/codex-usage-stdout

Conversation

@orienw

@orienw orienw commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Codex 0.156 can send notifications and an RPC reply in the same stdout chunk. readline() buffers the reply, but the next select() only checks the pipe. The collector then times out on account/read and the panel shows an "account/read" error despite having received the answer.

Read raw bytes into a shared buffer and consume complete lines before waiting on the pipe. Partial replies still respect the request timeout.

Validation:

  • Added regression coverage for batched notifications and replies, split replies, stalled partial replies, and EOF. The batched case fails on the original collector.
  • Collector, usage-update, and agents-panel tests pass.
  • Live Codex 0.156: the installed collector failed 5 of 6 runs; the patched collector passed all 6 in 0.80–1.09 seconds.
  • ./test/all: CLI suite passes. Seven shell test files fail, all reproduced on clean upstream: config, kernel-headers-migration, omarchy-kernel-migration, runtime-smoke, screenshot-sanity, snapper, and unowned-system-paths. Three other files report skipped checks.

@llstrk

llstrk commented Sep 23, 2026

Copy link
Copy Markdown

Verified: No actionable defect found in the scoped review of af2da3b6b4518. The collector now consumes complete buffered messages before waiting for more pipe data, fixing the notification-plus-reply timeout reproduced with a synthetic Codex backend.

Synthetic RPC input Base collector PR collector
Notifications and reply in one write account/read timeout Expected tier and weekly limit returned
Reply split across writes Expected result returned Expected result returned
Incomplete reply, then stall or EOF Explicit limits error Explicit limits error; no indefinite wait

The batched case in the added test targets the reproduced failure: buffered readline() can retain the reply while select() waits on an empty pipe. The explicit byte buffer removes that mismatch.


Review information

Test scope: Source inspection and isolated execution of both collector revisions with synthetic backends. No live Codex/account access, panel UI validation, or full-suite rerun; the reported live Codex 0.156 results were not independently verified.

Community review: Independent automated community review, unaffiliated with the Omarchy team, intended to help prepare PRs for their review.

Automated AI review: Astra Medium performed initial inspection and synthesis; Opus 5.5 High and GPT 6 Sol Xhigh completed independent technical reviews. Claims were checked against source and retained targeted evidence, followed by a separate Astra Medium editorial check.

@sanjyay

sanjyay commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Tested in the Omarchy VM (Omarchy 4.0.4 / Python 3.14):

  • Bug reproduction verified: Executed the test against the unpatched omarchy-agent-usage-codex collector. When Codex sends notifications batched with the RPC reply in a single stdout chunk, select.select() on the pipe stalled waiting for kernel bytes while the reply was buffered in Python's stream, timing out and returning Codex limits unavailable account/read.
  • Fix verified: With bufsize=0, explicit bytearray buffering, and draining complete lines before calling select(), the patched collector handled batched notifications and replies, split payloads, stalled partial replies, and EOF without timing out or hanging.
  • Test suite:
    • bash test/shell.d/agent-usage-codex-limits-test.sh — passed all 4 test modes (batched, split, stalled, eof).
    • bash test/shell.d/agent-usage-codex-scanner-test.sh — passed all 21 assertions.
    • bash test/shell.d/agent-usage-update-test.sh — passed all 7 assertions.

Works as expected!

viganogabriele pushed a commit to viganogabriele/agent-usage-plus that referenced this pull request Sep 25, 2026
)

Codex CLI 0.156 writes notifications (remoteControl/status/changed,
account/updated) and the RPC reply in a single stdout chunk. Omarchy's
packaged collector selects on the pipe, then calls readline(), which pulls
the whole chunk into Python's buffer and returns only the first line. The
next select() waits on an empty pipe while the reply sits in that buffer,
so account/read times out and the panel shows "Codex limits unavailable"
with "account/read" underneath, for a healthy ChatGPT login.

Add a third patch to omarchy-agent-usage-codex-compat that reads raw bytes
into a buffer kept on the process and drains complete lines before waiting
on the pipe. It mirrors the upstream fix in omacom/omarchy#12979 (issue
#10143) but replaces only rpc_request's body, so the call sites stay
untouched, and like the other patches it is skipped once Omarchy ships the
fix.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@NOirBRight

Copy link
Copy Markdown

Verified locally on Omarchy 4.0.4-1 / codex-cli 0.157.1 (Linux x86_64).

Applied the transport changes from this PR to a user-owned copy of the installed collector, leaving the original 4-second account/limits deadlines unchanged:

  • Six consecutive live --limits-only probes succeeded and returned the weekly limit (1.63–2.39s end-to-end).
  • Isolated transport probes passed for batched notifications/replies, split replies, stalled partial replies, and EOF.
  • Routed my existing user-owned agents widget through the patched collector and verified the actual panel displays the plan and weekly limit instead of account/read.

Before the fix, instrumentation at the original timeout boundary proved that the matching successful response (id 2) was already in the Python text buffer while the kernel fd was not readable and the server was still running. Details: #13080 (comment)

Scope: local transport probes, live collector, and panel verification; I did not run this PR’s full repository test suite.

@NOirBRight

Copy link
Copy Markdown

Follow-up to my earlier verification: the reader fix is present and matches this PR, but the unchanged 4s deadlines still cause independent intermittent failures.

Captured an actual widget refresh after applying the reader fix. Two local patched collector processes logged successful account/read responses in 2.048s and 3.050s, then timed out on account/rateLimits/read at 4.005s and 4.004s. At each timeout the explicit receive buffer was empty and the app-server was still running. The generated failure timestamps matched the widget’s codex.json, ruling out an unpatched writer for this captured failure.

With an extended observation deadline, real successful account/read responses arrived at 7.769s and 7.783s, and one rate-limits response arrived at 4.215s. Thus account/read, not only rateLimits/read, can legitimately exceed the current 4s budget here. This is distinct from the buffered-reader race fixed by this PR and consistent with #10143.

A user-local follow-up allowing 20s for both account RPCs passed delayed-account/delayed-limits collector regression cases and four real widget refreshes (8.25–10.44s total). No retries or stale-data fallback were added. This does not invalidate the reader fix; it narrows my earlier successful live-test claim, which did not exercise the later slow-response condition.

@markrotter

Copy link
Copy Markdown

Verified on omarchy 4.0.4-1 + codex-cli 0.157.1, and this PR's pending-buffer approach fixes it. Adding reproduction data in case it helps.

The race is deterministic, not intermittent, on 4.0.4. The app-server answers initialize and then pushes remoteControl/status/changed and account/updated in the same write. readline() returns the first line and strands the other two in Python's internal buffer, so the select() in the next rpc_request watches an already-drained fd and blocks until its deadline. Every request after the first initialize therefore times out.

Because the collector reports str(exc) as authHelpText and Panel.qml renders that string in its red status box, the user-visible symptom is the literal text account/read in the agents panel, with no Codex rate limits. That also makes this look like a credentials problem when it is not — codex login status is fine, and the same session answers both requests in ~1s when read without select().

Reproduce against the installed collector:

# verbatim rpc_request() from bin/omarchy-agent-usage-codex @ 4.0.4
p = subprocess.Popen(["codex","-s","read-only","-a","on-request","app-server"],
                     stdin=subprocess.PIPE, stdout=subprocess.PIPE,
                     stderr=subprocess.DEVNULL, text=True)
rpc_request(p, 1, "initialize", {"clientInfo": {"name": "t", "version": "1"}}, timeout=8)
# reads: initialize result, remoteControl/status/changed, account/updated
rpc_request(p, 2, "account/read", timeout=4)
# TimeoutError: account/read

Sending both requests up front, or reading the stream without select(), returns in 0.8s / 1.3s:

{"account": {"type": "chatgpt", "email": "...", "planType": "pro"}}
{"rateLimits": {"primary": {"usedPercent": 1, "windowDurationMins": 10080, "resetsAt": ...}}}

With the reader fixed, --limits-only goes from a 4s+ timeout to 2.3s and populates tierLabel: "pro" and limits: [{"label": "Weekly (7-day)", "percent": 0.01, ...}], with local session stats unchanged.

One note on the 4s deadlines in this PR: with the reader fixed, account/read and account/rateLimits/read are not close to the limit (they answered in ~0.5s and ~1.3s on a warm login), so the existing timeouts are fine as-is. The separate text=True + partial-write handling is worth keeping, since bufsize=0 turns stdin into a raw FileIO where a short write would otherwise silently truncate a request.

Also relevant: #13059 fixes the other half of the report — the bare method name shown to the user. Worth landing together, since #13059's clearer message is what makes this diagnosable at all.

@dhh

dhh commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Thanks! The same fix is coming in through #14049 (from #13703, which reads replies from the raw fd on top of #13733's request order). Closing in favor of that.

@dhh dhh closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants