Skip to content

Optimize Codex usage session scanning - #12336

Open
nick1udwig wants to merge 1 commit into
omacom:quattrofrom
nick1udwig:agent-usage-codex-incremental-scan
Open

nick1udwig wants to merge 1 commit into
omacom:quattrofrom
nick1udwig:agent-usage-codex-incremental-scan

Conversation

@nick1udwig

Copy link
Copy Markdown

Problem

omarchy-agent-update takes ~15s on my machine running at 100% CPU.

Solution

Cache what we've already seen.

Notes

Went from ~15s -> ~0.01s for a ~/.codex scan of ~1400 files and ~4.6GB. Cache for these codex files is ~600KB.

Very similar to #8313 but for codex. Implementation details differ.

@llstrk

llstrk commented Sep 25, 2026

Copy link
Copy Markdown

Automated AI review

Community review: Independent automated community review, unaffiliated with the Omarchy team, intended to help prepare PRs for their review.

Verified: On synthetic Codex sessions, the incremental scan produced the same collector output as the pre-PR full scan for every Codex-style write pattern tested, and a warm run opened no unchanged transcript. Two low-severity issues remain: an error-handling regression that can stop the whole collector, and a lost resume point for files that change mid-read.

Case (synthetic sessions) Head vs. pre-PR output
Cold scan, warm scan, appends that switch models, half-written last line identical
New, deleted and archived files, timezone change, truncation, --force identical
Rewrite inside the first 4 KiB plus growth, two collectors started together identical
Writer appending during 35 back-to-back scans final output identical; interim results never above the true count
Warm scan with nothing changed identical, 0 transcripts opened

A same-inode edit between the two 4 KiB anchor windows, followed by growth, is treated as an append and not detected (--force restores the correct totals). Codex's append-only writer does not produce that pattern.

Verified: The new agent-usage-codex-incremental-test.sh passes on this head and fails against the pre-PR script. The existing agent-usage-codex-scanner-test.sh still passes. The Codex CLI rollout recorder (0.154.0 and 0.157.0) opens rollouts for append and writes each record plus its newline in one write_all. In the default configuration it does not rewrite rollouts in place, so the append-only assumption holds.

Narrower exception handling can stop the whole collector

The pre-PR loop caught any exception per file and skipped that file. The new code catches only ValueError/TypeError/AttributeError per line (L364) and OSError per file (L411). Any other exception now escapes. For example, a numeric "timestamp": 1e300 reaches datetime.fromtimestamp in local_day outside its try (L48) and raises OverflowError:

one session file containing a line with "timestamp": 1e300, run with --force

pre-PR: skip that file ............................. exit 0, JSON record printed
head:   OverflowError -> "cache unavailable; scanning directly"
        -> direct scan hits the same line -> exit 1, empty stdout

Impact: One malformed session file makes the collector return no usage data at all, and the stderr warning blames the cache. Codex itself writes string timestamps, so the trigger is a corrupted or foreign file under sessions/ or archived_sessions/.

Suggested change: Keep a broad per-file except Exception: continue around read_native_session and the row loop, as before, or move the numeric branch of local_day inside its try.

A file that changes during its read loses its resume point

When the final fstat differs from the opening signature, the new snapshot is correctly not persisted (L368-L369). However, only records with a signature are written to the new cache (L409-L410), so the previous valid record for that file is dropped too:

run N:   valid cached record -> append read -> writer appends meanwhile
         -> snapshot discarded AND previous record discarded
run N+1: no cached record -> full re-read from byte 0

Impact: Totals stay correct, but a large session that is still being written can be re-read in full on every refresh instead of resuming. With an aggressive synthetic writer, the active file was never cached in 35 of 35 overlapping runs. How often this happens with real Codex write rates was not measured.

Suggested change: When the read took the append path, keep previous in updated if the new snapshot cannot be persisted, so the next run resumes from the last good offset.

Unterminated final record: A complete, valid final record with no trailing newline is counted by the pre-PR script but never by this head, even with --force or after archiving (L338). It is counted once more bytes are appended. Codex writes the newline in the same write and repairs a missing one when it reopens a session, so this needs an interrupted write that ends exactly after } in a session that is never resumed. Counting a tail that parses as a complete object, without advancing the offset, would restore parity if wanted.


Review information

Test scope: Reviewed head 0621e2fa. Runtime checks used sandboxed synthetic Codex session files only, comparing this head with the merge-base script. Real Codex usage data and real-world timing were not tested, so the reported 15 s to 0.01 s improvement was not independently reproduced. On 120 synthetic files, the native cache was about 320 bytes per file. Codex recorder behavior comes from source reading of releases 0.154.0 and 0.157.0. The version the author tested is not stated.

AI process: Opus 5.5 Medium coordination and synthesis, independent Opus 5.5 Xhigh and GPT 6 Sol Xhigh technical assessments with targeted follow-up questions, Opus 5.5 Medium editorial check.

Opt out: To stop receiving these reviews, reply to this comment saying so.

@omarchybot

Copy link
Copy Markdown
Collaborator

#14049 has merged into quattro and reuses unchanged native rollout totals, so that portion overlaps. This PR additionally resumes reading at the prior byte offset when a rollout grows; the landed scanner rereads changed files in full. Keeping this open for that distinct append-only optimization. Rebase onto the current scanner and preserve its provider filtering and error handling.

Closure scope checked by GPT-6 in Codex and Codex Medium against the reports and landed source. The second opinion agreed; independence is not guaranteed. No runtime tests were rerun for this reconciliation.

@nick1udwig
nick1udwig force-pushed the agent-usage-codex-incremental-scan branch from 0621e2f to f295191 Compare October 9, 2026 18:20
@nick1udwig

Copy link
Copy Markdown
Author

@omarchybot updated in f295191

@greptile-apps

greptile-apps Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium impact] The PR appears safe to merge; no actionable issues were found.

Summary

The PR resumes scans of growing Codex session files instead of reading their full history again.

  • Codex usage scans resume at the end of growing session files.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart TD
  A[Find recent session files] --> B{Saved signature matches?}
  B -->|Yes| C[Reuse saved totals]
  B -->|No| D{File grew and saved samples match?}
  D -->|Yes| E[Restore state and read new lines]
  D -->|No| F[Read from the beginning]
  E --> G{Read stayed stable and complete?}
  F --> G
  G -->|Yes| H[Save new totals and offset]
  G -->|No| I[Keep the last safe cache record]
  C --> J[Merge usage totals]
  H --> J
  I --> J
Loading

Reviews (1) · Last reviewed commit: "Optimize Codex usage session scanning" · Reviewed by Greptile

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants