Skip to content

fix(agents): give Codex rateLimits RPC more time and clearer timeout text (#12880) - #12984

Closed
kvnloo wants to merge 0 commit into
omacom:quattrofrom
kvnloo:fix/12880-codex-ratelimits-timeout
Closed

kvnloo wants to merge 0 commit into
omacom:quattrofrom
kvnloo:fix/12880-codex-ratelimits-timeout

Conversation

@kvnloo

@kvnloo kvnloo commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

What

Raise account/read and account/rateLimits/read RPC timeouts from 4s to 8s (same budget as initialize).

When rpc_request times out, raise TimeoutError with Codex RPC timed out waiting for {method} instead of the bare method name.

Why

Under load the Codex app-server often needs more than 4s for account/rateLimits/read. The collector stored str(TimeoutError) as authHelpText, so the panel showed account/rateLimits/read and looked like an auth/unavailable failure (#12880).

Tests

bash test/shell.d/agent-usage-codex-scanner-test.sh

All assertions passed, including source pins for the 8s timeouts / clear TimeoutError text and a short-timeout unit check that the message names the method and says timed out.

Pinned tip: d3cfd53b997f8bdcf776b8db68bf0d735e7a065d

Fixes #12880

@llstrk

llstrk commented Sep 23, 2026

Copy link
Copy Markdown

No important actionable issue found in the inspected scope at 3996942f.

Verified by source inspection: both account RPCs now receive the same eight-second budget as initialization, and the descriptive exception reaches the panel through authHelpText.

RPC setting Before This PR
initialize 8 seconds 8 seconds
account/read 4 seconds 8 seconds
account/rateLimits/read 4 seconds 8 seconds
Timeout text Bare method name Codex RPC timed out waiting for {method}

These are per-call budgets, not an eight-second total deadline. The inspected updater and panel impose no shorter timeout that would override them.


Review information

Test scope: Source review of the collector, added tests, updater and panel integration. No runtime result is relied on here; the repository test suite, live panel and authenticated Codex behavior under load were not validated. Whether eight seconds is sufficient under real load remains unmeasured.

Community review: Independent automated community review, unaffiliated with the Omarchy team, intended to help prepare PRs for their review.

Automated AI review: Astra Medium initial inspection and synthesis, independent Opus 5.5 High and GPT 6 Sol Xhigh technical reviews, with claims checked against the pinned source.

@sanjyay

sanjyay commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Tested in the Omarchy VM (Omarchy 4.0.4):

The clearer TimeoutError(f"Codex RPC timed out waiting for {method}") text works nicely and prevents authHelpText from displaying just the raw method name.

However, testing the 8s timeout under real/simulated Codex stdout traffic revealed that increasing the timeout budget does not resolve the root timeout:

Issue observed in VM

When Codex sends notifications or messages batched in stdout chunks (common during startup or under load), Python's TextIOWrapper (text=True on proc.stdout) buffers the stream in user space via readline(). On subsequent requests, select.select([proc.stdout], [], [], 0.25) checks the underlying OS file descriptor, which has no remaining kernel bytes, causing select() to stall until the deadline.

In testing:

Suggested fix

The timeout is primarily caused by this I/O buffer stall rather than Codex taking >4s to respond. PR #12979 addresses this root cause by switching to unbuffered binary reads (bufsize=0, os.read(), and draining a bytearray buffer before calling select()). Combining PR #12984's clearer error messaging with PR #12979's unbuffered stream draining would provide both the root fix and better error diagnostics.

@kvnloo kvnloo closed this Sep 23, 2026
@kvnloo
kvnloo force-pushed the fix/12880-codex-ratelimits-timeout branch from 3996942 to c737c49 Compare September 23, 2026 18:15
@kvnloo

kvnloo commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

Withdrawn in favor of #12979.

Runtime testing showed that increasing the Codex RPC timeout from 4s to 8s does not fix the failure; it only makes the same buffered-I/O stall take longer. #12979 addresses the reproduced root cause by draining buffered stdout before waiting on the file descriptor.

I reset this branch to current "quattro".

The clearer timeout diagnostic from this PR may still be useful as a small follow-up after the I/O fix lands, but it should not compete with the root-cause patch.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

omarchy-agent-usage-codex: account/rateLimits/read RPC timeout (4s) misreported as unavailable

3 participants