Skip to content

io patch: guard fiber-scheduler fd dispatches on embedded fds - #140

Merged
ronaldtse merged 1 commit into
mainfrom
patch/io-scheduler-guards
Oct 8, 2026
Merged

ronaldtse merged 1 commit into
mainfrom
patch/io-scheduler-guards

Conversation

@ronaldtse

Copy link
Copy Markdown
Contributor

Closes #138.

Root-caused downstream in tamatebako/tebako#252 (full repro + evidence chain there): with a fiber scheduler installed, io.c's read/write dispatches hand the fd to the scheduler hook — Async::Scheduler#io_read → io-event's C extension — which issues raw libc fcntl/read on it. A VFS-backed File holds a tagged virtual fd (fd | 0x40000000) that only works because this patch macro-wraps the fd verbs into tfs_* dispatchers inside ruby core's translation units; the extension's calls escape that, the packaged runtime has no interposition tier for them, and the kernel answers EBADF (read(1073742315) — the exact fd on the issue).

Fix

Guard every fiber-scheduler fd dispatch in io_c_tebako_includes.patch with tebako_fd_is_embedded(...) — the same idiom the patch already uses for copy_file_range/sendfile. Embedded fds are served synchronously by the in-process engine and never block, so the guard skips the scheduler (the direct wrapped path serves the op — no concurrency is lost), and the wait funnel answers "ready" immediately instead of dispatching io_wait.

Per line (ruby 3.4.8 io.c line numbers cited on the issue; per-line equivalents found by hand, context verified byte-identical across every release of each line):

  • 3.2 (5 sites): rb_io_read_memory, rb_io_write_memory, rb_writev_internal, io_read_memory_call (the readpartial/read_nonblock dispatch — same io_read_memory hook, needed to cover the whole non-blocking read surface), and the rb_io_maybe_wait EAGAIN funnel (covers rb_io_maybe_wait_readable/_writable and every direct caller; embedded fd ⇒ return the requested events immediately). No pread/pwrite dispatch exists on this line (added upstream in 3.3) — nothing to guard there.
  • 3.3 (7 sites): same as 3.2 plus pread_internal_call / internal_pwrite_func (guard on arg->fd).
  • 3.4 (7 sites): same as 3.3 (the issue's line: rb_io_pread/rb_io_pwrite dispatches live in the *_internal_call/*_func helpers). io_close needs no guard here — its scheduler dispatch is commented out upstream (io.c:5570).
  • 4.0 (9 sites): same as 3.4 plus two dispatches that are live upstream on this line: io_flush_buffer_async (buffered-write flush → io_write_memory hook) and io_close (rb_fiber_scheduler_io_close(scheduler, RB_INT2NUM(fd)) — the 4.x re-check the issue asked for: unguarded, the hook's raw close(2) dies EBADF and a truthy result would skip the wrapped close, leaking the memfs fd).

One extra site beyond the issue's literal list is guarded everywhere: io_read_memory_call (the scheduler dispatch behind IO#readpartial/IO#read_nonblock). It is the same io_read_memory hook escaping the same way; leaving it open would keep readpartial broken under Async. Recorded here per the issue's "guard every fiber-scheduler fd dispatch".

Still out of scope (adjacent gaps recorded on the issue, unchanged): IO#wait/rb_io_wait and rb_io_wait_readable/_writable on raw fds (selector registration — unreachable on the #252 read path since memfs reads never EAGAIN), IO.select/rb_fiber_scheduler_io_selectv on tagged fds.

3.1 line

Checked per the task: 3.1's io.c does carry the same read/write/writev/io_read_memory_call scheduler dispatches and the same rb_io_maybe_wait funnel (no pread/pwrite hooks). Per the task disposition the 3.1 line keeps its current surface in this PR; if 3.1 runtimes ever matter for Async workloads, the same guard set applies mechanically (context hashes match 3.2's shape).

Verification

  • git apply --check of each regenerated patch against every release of its line (35/35 clean: 3.2.4–3.2.11, 3.3.3–3.3.12, 3.4.1–3.4.10, 4.0.0–4.0.7); applied result diffed against the intended tree (byte-identical). Patch bodies regenerated with git diff --no-index from the pristine tarballs — the pre-existing hunks (include block, copy_file_range, sendfile) are byte-identical modulo @@ offsets.
  • tools/validate_manifests — OK.
  • tools/lint for all 37 versions in versions.yml (3.1 control line included, its patch untouched and byte-identical) — all clean.
  • bundle exec rspec — 121 examples, 0 failures (no selection-metadata change, so no new specs).
  • tools/compile_smoke (io.o + all patched TUs) per changed line's tip version: 3.2.11 / 3.3.12 / 3.4.10 / 4.0.7 on darwin (host) — green; msys x86_64-w64-mingw32 cross legs — green. CI runs the full diff-aware smoke matrix (incl. darwin on macos-latest and the msys cross leg) on this PR via lint-patches.
  • Functional Async repro: not run locally. The escape needs a packaged runtime (the in-process memfs engine + a fiber scheduler); building one is the runtime factory's job and exceeds this repo's harness (compile-smoke compiles against stub c_api headers, it does not link a runnable tebako ruby). The packaged fixture arms that close Check Zeitwerk compatibility tebako#252 live in tebako's suite and gate the next tebako-runtime-ruby build carrying this patch — tracked downstream. The guard logic itself is verifiable by inspection against the pinned mechanism: guarded sites can no longer hand a tagged fd to the scheduler, and the synchronous fall-through uses the already-wrapped read/pread/pwrite/writev paths (memfs reads never EAGAIN, so rb_io_blocking_region_wait's retry path is never entered).

Patch mechanics

Diffs regenerated with git diff --no-index against the pristine 3.2.11/3.3.12/3.4.10/4.0.7 io.c (each line's tip); @@ trailing context stripped to match repo style. Guards added as comment + extended if condition at each dispatch site — no manifest/selection changes (whole-line features, unchanged names), so selection for every previously-selected version is byte-identical in shape and the 3.1 folder is untouched.

With a fiber scheduler installed, io.c hands the fd to the scheduler
hook (Async::Scheduler -> io-event's C extension), which issues raw
libc fcntl/read on it -- EBADF for an embedded (memfs) fd, a flagged
virtual descriptor the kernel does not know (tamatebako/tebako#252).

Guard every fiber-scheduler fd dispatch in io_c_tebako_includes.patch
with tebako_fd_is_embedded() -- the patch's existing
copy_file_range/sendfile idiom. Embedded fds are served synchronously
by the in-process engine and never block, so the dispatch is skipped
(the direct wrapped path serves the op) and the rb_io_maybe_wait
funnel answers "ready" immediately instead of waiting.

Sites per line (context verified byte-identical across every release
of each line; patches regenerated with git diff --no-index):

- 3.2 (5): rb_io_read_memory, rb_io_write_memory, rb_writev_internal,
  io_read_memory_call (the readpartial/read_nonblock dispatch -- same
  io_read_memory hook), rb_io_maybe_wait. No pread/pwrite dispatch
  exists upstream on this line.
- 3.3/3.4 (7): the 3.2 set plus the pread/pwrite dispatches
  (pread_internal_call / internal_pwrite_func, guard on arg->fd).
- 4.0 (9): the 3.4 set plus the two dispatches that are live upstream
  on this line: io_flush_buffer_async and io_close (its hook's raw
  close(2) would die EBADF and a truthy result would skip the wrapped
  close, leaking the memfs fd).

3.1 keeps its current surface per the issue disposition; its io.c
carries the same read/write/writev/maybe_wait dispatches (noted on
the PR).

Verified: git apply --check clean on all 35 releases of the touched
lines; tools/lint green on all 37 versions (3.1 control untouched);
tools/validate_manifests OK; rspec 121 examples 0 failures;
compile_smoke green on each line's tip (darwin + msys cross).

Closes #138
@ronaldtse
ronaldtse merged commit 0275c07 into main Oct 8, 2026
60 checks passed
@ronaldtse
ronaldtse deleted the patch/io-scheduler-guards branch October 8, 2026 14:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

io patch: fiber-scheduler fd dispatches escape the tfs_* wraps (Async/io-event EBADF on VFS fds) Check Zeitwerk compatibility

2 participants