Skip to content

feat: forked python process - #706

Open
doc-han wants to merge 4 commits into
mainfrom
forked-python-process
Open

doc-han wants to merge 4 commits into
mainfrom
forked-python-process

Conversation

@doc-han

@doc-han doc-han commented Oct 2, 2026 •

Copy link
Copy Markdown

Short Description

Adds an opt-in warm fork-server for Python services: a long-lived master pays the startup cost (poetry run + heavy imports) once and forks a fresh child per request, cutting warm per-request latency ~25× while
keeping the per-request process isolation. This is gated behind APOLLO_FORK_SERVER=true; the spawn path is unchanged when off.

Fixes #

Implementation Details

Today every request spawns poetry run python services/entry.py, paying ~210 ms
of poetry run + ~450 ms of heavy imports (sentry/opentelemetry/langfuse/anthropic)
before any work. This PR pays that once in a warm master and forks per request.

  • services/entry.fork.py (new) — the master. Preloads entry.py's heavy
    libraries once with no init, so it stays single-threaded and fork-safe, then
    os.fork()s a fresh child per request over a unix socket. Each child imports
    entry.py after forking
    , so entry.py's own init (OTel/Langfuse/Sentry) runs
    in the child where forking is safe — this is what lets entry.py stay
    untouched
    . The child dup2()s the socket onto stdout/stderr so service
    logging streams unchanged, emits PID:/EXIT: control lines, and returns the
    result via the existing output-file contract.
  • platform/src/bridge.fork.ts (new) — all Node-side fork code (warm-master
    singleton, unix-socket client, runForked). Same output-file + error-mapping
    contract as the spawn path.
  • platform/src/bridge.ts — 3-line change: import, flag, one if branch.
    The spawn path is otherwise untouched.
  • .env.example — documents APOLLO_FORK_SERVER and, for macOS,
    OBJC_DISABLE_INITIALIZE_FORK_SAFETY=YES (a forked child touching an Obj-C
    framework otherwise aborts; no-op on Linux, so prod needs nothing).

AI Usage

Please disclose whether you've used AI in this work (it's cool, we just want to
know!):

  • Yes, I have used AI
  • No, I have not used AI

You can read more details in our
Responsible AI Policy

@josephjclark

josephjclark commented Oct 2, 2026 •

Copy link
Copy Markdown
Collaborator

This PR just sets up the forking, but otherwise leaves all IPC just the same as prod

We could consider deploying this directly to make prod apollo run faster

TODO:

  • Joe to carefully test
  • Including the CLI
  • And including internal token stuff

Consider:

  • refactoring to preserve both child process strategies
  • Spitting all child-process calls out into two files/folders
  • On MacOS, if that env var isn't set, fallback to the spawn appraoch and log a warning

@doc-han
doc-han requested a review from josephjclark October 7, 2026 11:47
@doc-han
doc-han marked this pull request as ready for review October 7, 2026 11:47
Comment thread services/entry.fork.py
# Warm entry.py's heavy deps so forked children inherit them (copy-on-write).
# Mirrors entry.py's imports; a missing one is just imported per-child - slower,
# never wrong.
PRELOAD = (

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'd like to pull this preload list out of here and into a yaml file, with some comments to explain what it is and what it's doing

Also, since 95% of calls will be through global_chat, can we add its heaviest dependencies here? Claude suggests Pinecone and openai, but maybe there are others.

@josephjclark

Copy link
Copy Markdown
Collaborator

I ran some review with claude which is worth checking up on:

The following issues mean the forks are not safe to deploy to prod:

  1. Sentry events get dropped. The child exits with os._exit (entry.fork.py:83), which skips Sentry's flush. The fix is sentry_sdk.flush(timeout=2) plus langfuse.flush() before exiting.
  2. A master that dies before it's ready hangs the request forever. The exit handler in bridge.fork.ts:68-73 never rejects the pending promise, and the 60-second timeout never kills the process. It triggers on a bad preload import or a failed bind.
  3. The Sentry service tag is lost. entry.main() sets it (entry.py:196) and the fork child skips it. It's a one-line fix.
  4. Python tracebacks vanish from server logs. Non-log stderr isn't forwarded the way bridge.ts:156-158 does it.

I think that 4th one relates to calls to apollo through the CLI - a super cool feature which has never been used. I think we can afford to drop that stuff BUT Id' rather formally remove it from the server and CLI than just leave it broken. If there's an easy fix, I'd like to adopt it for now.

Also some suggested manual tests:

  1. Sentry: force a service exception and an ApolloError, and check that both reach Sentry with the service tag. Compare against flag off.
  2. Broken preload: put a bad module in PRELOAD. The request should fail fast, not hang.
  3. Master crash: kill -9 the master mid-stream, then confirm the next request recovers.
  4. Client abort: drop an SSE client mid-stream and confirm the child is killed.
  5. Streaming parity: compare streaming output with flag off.
  6. Tracebacks: check that a Python traceback shows up in the server logs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants