Skip to content

Develop - #47

Merged
demonixis merged 106 commits into
mainfrom
develop
Oct 3, 2026
Merged

demonixis merged 106 commits into
mainfrom
develop

Conversation

@demonixis

Copy link
Copy Markdown
Owner

No description provided.

SBudarin and others added 30 commits June 24, 2026 23:29
Drive the immersive space from explicit user intent (wantsImmersiveSpace) instead of the raw connection state, so exiting via the Digital Crown returns to the menu instead of auto re-entering. Add an "Enter Immersive View" button (plus Disconnect) to the control window for re-entry while still connected.

Recreate the single-use ARKit session and data providers before each (re)entry so head tracking resumes on reconnect instead of running a permanently-stopped provider.
Add H.264/H.265 codec negotiation across the runtime, clients, and Home frontends while keeping H.265 as the default for legacy clients.

Split swapchain handling by backend and add Linux OpenGL plus Windows D3D11/D3D12 runtime paths with FFmpeg readback support.

Fix CTS Metal mip-count semantics, propagate the Metal toolchain into the CTS sub-build, and update the related documentation and tests.
# Conflicts:
#	AGENTS.md
#	README.md
Advertise 10-bit decode support from visionOS and enable VideoToolbox Main10 only for capable H.265 clients. Preserve the 8-bit H.264 and legacy-client paths, handle 10-bit decoder surfaces in the Metal renderer, and retain the setting across stream reconfiguration.

Expose the option in SwiftUI Home and Qt Home, add protocol and configuration coverage, and document the behavior and visionOS UI workflow.
Define VideoToolbox streams as limited-range BT.709 SDR by setting explicit primaries, transfer-function, and YCbCr-matrix metadata.

Replace the full-range BT.601 shader conversion with BT.709 conversion using exact normalized limited-range constants for both 8-bit and 10-bit decoder surfaces. This restores black levels and color balance without changing encoded bandwidth.

Document the streaming color contract and its client requirements.
The visionOS client requests a 10-bit HEVC decode surface for every H.265
stream, but the server only enables Main10 when it is configured for it *and*
the client advertised the capability. The default configuration therefore
streams 8-bit while the client still asks the decoder for a 10-bit output
surface. If VideoToolbox refuses a 10-bit surface for an 8-bit stream,
VTDecompressionSessionCreate fails outright and nothing decodes — a black
screen on the default path, not an edge case.

VideoDecoder now tries the preferred 10-bit surface first and falls back to
8-bit if session creation fails, logging which surface it settled on. The
renderer already selects its color conversion from the buffer's *actual*
pixel format, so whichever surface VideoToolbox hands back displays correctly;
this only removes the hard-failure path.

ImmersiveRenderer now always binds the color-params buffer (fragment buffer
index 1), defaulting to the 8-bit constants, so the fragment shader never
reads an unbound buffer on frames that draw before the first decoded pixel
buffer arrives (previously it was bound only inside the pixel-buffer branch
while the draw ran unconditionally).
Set kVTDecompressionPropertyKey_RealTime on every decompression session so
VideoToolbox prioritizes latency over throughput and does not batch or hold
frames. It is never combined with MaximizePowerEfficiency (undefined behavior).

Prefer the hardware decoder via the decoder specification
(kVTVideoDecoderSpecification_EnableHardwareAcceleratedVideoDecoder) so a
real-time stream is never paced by a software decoder. Enable — not Require —
so it can never hard-fail session creation; hardware HEVC/H.264 decode is
always present on Apple silicon, and this composes with the existing
10-bit->8-bit output-format fallback.

Split NAL units in place over the frame buffer via withUnsafeBytes instead of
first copying the whole frame into a [UInt8] array. Each NAL is copied out once
into an owned, 0-based Data — required so callers can index nal[0] and so
parameter sets can be retained across frames — which removes one full-frame
heap copy and the per-byte bounds checking on the decode hot path.

Document the decode changes, on-device verification, and the remaining latency
roadmap in docs/vision-pro-latency.md. Record the decode-path invariant in
AGENTS.md and add CHANGES.md entries for the low-latency decode work and for
the 10-bit->8-bit decode-surface fallback fix that landed without one.
Async-decode-off, in-flight-frame count, fixed compositor depth, and an
adaptive jitter buffer are called out as isolated experiments to measure
separately rather than bundled here.
The runtime advertised a fixed 1512x1680 per-eye render resolution to the app
regardless of the target headset, so Vision Pro streamed at a Quest-tuned
resolution. The render resolution is fixed when the app queries the view
configuration, which happens before any client connects, so it cannot be
chosen automatically per connected device — make it a server-config choice.

Add `render_device` ("quest2", "quest3", "avp") to select the per-eye render
resolution the runtime advertises via recommendedImageRect (1440x1584 /
1512x1680 / 3024x3360). The fallback-FOV aspect in InputManager derives from
the selected device. StartStreamingIfNeeded now sizes the streaming server
from the same preset (it was hardcoded to 1512x1680, which pinned the encoder
to that resolution no matter what the app rendered), so the encode matches the
render target. Use the existing `resolution_scale` to trim how much of that is
encoded and streamed (also what ABR drives); this preset only changes the
render target. Expose the device in SwiftUI Home and Qt Home. The default
(quest3) reproduces the previous 1512x1680.
The runtime already foveates the encoded stream and announces the layout, but
the visionOS client never advertised foveated-encoding support, so the server
fell back to non-foveated video and `foveated_encoding_preset` had no effect on
Vision Pro.

The client now advertises CLIENT_CAPABILITY_FOVEATED_ENCODING alongside 10-bit,
computes the inverse-AADT parameters from the server announce (center size and
shift, edge ratio, and a derived eye-size ratio that matches the Quest client's
active-ratio math), and hands them to the fragment shader. The shader inverse-
warps each displayed eye-UV back to the encoded texel via the ALVR axis-aligned
mapping (bisection inverse of compressAxis) before the stereo split and the
existing BT.709 / 10-bit color conversion. enabled == 0 is an exact passthrough,
so non-foveated streams are unaffected.
Bitrate-limited, bilinearly-upscaled streamed video reads soft on the headset.
Add an optional post-decode sharpening pass configured server-side.

`client_sharpening` (0.0-1.0) is a new runtime config; the server carries it to
the client as a 0-100 percent in ServerAnnounce (repurposing the reserved slot,
so the wire layout is unchanged and ProtocolLayoutTests still pass). The
visionOS client reads it from the announce and applies an FSR-RCAS-style
contrast-adaptive sharpen in the fragment shader: the scene resolve is factored
into sampleSceneRGB() so the pass can resample the 4 neighbours through the
exact reprojection/foveation/color mapping, then unsharp against their average
clamped to the local min/max (no ringing). Strength 0 is an exact passthrough.

Expose the strength as a slider in SwiftUI Home and Qt Home.
A decode failure (from packet loss) leaves later inter frames referencing data
the decoder no longer has, so decoding them smears green/blocky corruption
forward until the next scheduled keyframe.

The decoder now enters a recovering state on any decode error: it drops VCL
slices and keeps nudging the server for a keyframe (the existing cooldown
rate-limits the requests) until an IRAP (H.265 NAL 16-23) or IDR (H.264 NAL 5)
arrives, then resumes. Packet loss now shows a brief clean freeze that
self-heals instead of propagating corruption. Recovery state is cleared on
codec switch / invalidate.
The CompositorServices frame clock (queryNextFrame → wait(optimalInputTime))
is the real pacer, so a third in-flight buffer only lets the CPU run an extra
frame ahead — about one frame (~11 ms at 90 Hz) of added latency with no
throughput gain for a video blit. Drop to 2 to shave that frame.
visionOS always reprojects the presented frame using drawable.deviceAnchor and
the depth buffer. Clearing depth to a near 2 m plane makes the compositor apply
positional parallax to what is really a flat 2D video frame, so head
translation warps the whole image ("swim"/zoom on forward motion).

Clear depth to a far (~1000 m) head-locked distance instead: the positional
reprojection then contributes ~nothing, leaving clean rotation-only warp that
matches the shader's reprojection (the same rotation-only approach ALVR and
Virtual Desktop use for streamed 2D frames). Depth is still valid, so the
device does not drop frames as black.
The target uses SWIFT_DEFAULT_ACTOR_ISOLATION = MainActor, which makes every
unannotated type main-actor-isolated. The decode/render helper state
(PixelBufferState, KeyframeRecoveryState, EyeProjectionState,
RenderPoseReprojector, FoveationState, PostFXState, FoveationShaderParams) is
accessed off the main actor — from the VideoToolbox decode callback, the
render actor, and UDP threads — so those calls raised "main actor-isolated in
a nonisolated context" warnings.

Mark the types `nonisolated` so their init/methods stay off the main actor.
They are already lock-guarded `@unchecked Sendable`, so runtime behavior is
unchanged; this only aligns the declared isolation with actual use and clears
the warnings (committable, unlike toggling the target's default isolation).
The tab bar used a manual Binding whose setter wrote the @published
`selectedTab` — SwiftUI can invoke that setter during a view update while
reconciling the TabView selection, which publishes a change mid-update and
logs "Publishing changes from within view updates is not allowed."

Bind the TabView directly to `$model.selectedTab`; SwiftUI's projected binding
defers the write correctly, so the warning goes away with identical behavior.
A new "Emulate controllers" toggle lets controller-only PCVR games run on
Vision Pro without physical spatial controllers.

HandGestureEmulator synthesizes a VR controller per hand from the 26-joint
skeleton: index pinch → analog trigger, middle/ring pinch → face buttons
(A/B right, X/Y left), little pinch (left) → menu, three-finger curl → grip,
wrist → 6DOF pose, all with pinch/curl hysteresis to avoid chatter. When an
Xbox-style gamepad is connected, compatibility mode instead takes the 6DOF
pose from the hand and the buttons/sticks/triggers from the gamepad, emulating
Meta Touch controllers; it takes priority over gesture emulation.

VisionTrackingManager only fills a hand's controller when no physical spatial
controller is present, applies a Meta/Touch orientation correction (the raw
hand frame points the opposite way), and gates both modes behind the toggle
(off by default). Emulated controllers ride the existing tracking packet, so
there are no protocol or runtime changes.
Clearing depth to ~infinity (to avoid forward-motion swim) also disabled the
compositor's positional reprojection, which is the only client-side
compensation for head translation. Up/down/sway then lagged the full
round-trip and felt sluggish.

Restore the ~2 m depth plane so the compositor applies translation parallax
again — translation feels responsive at the cost of some swim on content far
from 2 m (the standard streaming-VR tradeoff; the real fix is streaming a real
depth buffer for true 6DOF timewarp). Reverts the depth change from the earlier
"Clear depth to far" commit.
The hands+gamepad controller emulation is a power-user/compatibility feature,
not part of the normal viewing flow, so group it under a collapsible
"Developer" disclosure in the visionOS control window instead of sitting
alongside the everyday toggles.
The client inverted the server's AADT warp by 10-step bisection, evaluating
the forward compress_axis (~15 ops with branches) ten times per axis per
pixel — and the sharpening pass multiplied that by its five resolves. The
forward warp is piecewise (quadratic / linear / quadratic) and monotonic per
piece, so it has an exact closed-form inverse: the linear solution in the
center region and the stable small root of a quadratic in each edge region
(conjugate form, no cancellation).

Verified numerically against the server's compress_axis for all foveation
presets (light/medium/high, both axes) plus shifted-center variants: maximum
round-trip error ~2e-7 (fp32 machine precision) vs ~5e-4 for the bisection it
replaces — about 2000x more accurate at roughly a tenth of the ALU cost. The
now-unused forward compressAxis copy is removed from the shader.
The sharpening pass resolved the full scene five times per pixel — each tap
re-running the reprojection ray math, the foveation unwarp, and both texture
planes plus color conversion. With foveation on, that multiplied the warp math
by five for a post-process.

Restructure the fragment so the display-to-video mapping runs once per pixel
(displayToStereoUV), then apply contrast-adaptive sharpening on luma only, in
source space: the four neighbour taps step one video texel from the computed
UV instead of re-running the mapping, and chroma is sampled once (luma carries
virtually all perceived sharpness; sharpening chroma mostly adds ringing).
Marginal cost of sharpening drops from 4 full scene resolves to 4 single-plane
texture samples plus a few ALU ops, all in the same pass — no intermediate
texture, no extra render pass, no added latency.

Operating on raw luma code values is exact: the later video-range expansion is
affine and CAS (unsharp + min/max clamp) commutes with affine maps. Source-
space taps also sharpen at the content's real texel frequency rather than
display frequency, so upscaled streams don't get their bilinear staircase
amplified. Horizontal taps are clamped to the eye's half of the side-by-side
frame so sharpening never bleeds across the stereo seam. PostFXParams no
longer needs invResolution; the shader reads the tap size off the texture.
LayerRenderer.Drawable.View.tangents is deprecated since visionOS 2.0 in
favor of cp_drawable_compute_projection. Derive the per-view frustum tangents
from drawable.computeProjection instead: the matrix maps view space looking
down -Z with w_clip = -z (the depth-clear path already depends on this
convention working on device), so the NDC edge conditions invert to
tanLeft = (1 - P20)/P00, tanRight = (1 + P20)/P00, tanUp = (1 + P21)/P11,
tanDown = (1 - P21)/P11 — the same positive (left, right, up, down)
magnitudes view.tangents reported.

Both consumers switch to the shared helper: the reprojection shader's per-eye
frustum data and the FOV/IPD published to the runtime. The one-shot [ProjDiag]
log keeps its format, so the derived values can be compared on-device against
previously logged tangents. Clears the last three deprecation warnings.
Rotation was reprojected for the full render-to-display latency, but
translation was not: the shader warp was rotation-only, the echoed render-pose
position was parsed off the wire and discarded, and the compositor's
depth-based warp only covers the last milliseconds (deviceAnchor is the live
pose at presentation). Head translation therefore lagged the entire pipeline
latency — the "sluggish up/down/sway" feel.

Keep the render-pose position (VideoReceiver now delivers position and
orientation; RenderPoseReprojector stores the full pose) and extend the warp:
for a point assumed at the reprojection plane distance d along the current
ray, its render-eye position is rot*dir*d + delta, so the shader adds delta/d
to the rotated ray — planar parallax for the full render-to-now translation.
The per-eye delta includes the rotation-induced eye offset (IPD lever arm) and
is clamped to 0.5 m so a bad pose match cannot distort the warp; zero delta
remains an exact passthrough.

The plane distance is a shared constant with the depth-buffer clear, so the
shader warp and the compositor's positional pass always assume the same
scene depth. Parallax is exact at the plane and approximate away from it —
the same tradeoff as the compositor warp, but now covering the whole
pipeline latency instead of the final milliseconds.
The runtime's bounded pose prediction prefers client-reported velocity and
only falls back to finite-differencing the received poses — a noisy signal
off jittery UDP arrival that keeps the usable prediction conservative. The
visionOS client never filled headLinearVelocity/headAngularVelocity, so it
always got the fallback.

Measure both velocities on the tracking queue by differencing consecutive
ARKit device-anchor samples taken on the steady 90 Hz timer, with a light
(~2-sample) EMA to strip per-sample noise. Angular velocity is world-frame
from the pre-multiply delta (q_now = dq * q_prev), matching the convention
PredictOrientationFromVelocity applies on the server. Velocities reset to
zero when tracking is lost and restart cleanly on reacquisition.

With clean velocity in, the server's existing prediction path renders frames
closer to where the head actually is, cutting the translation error the
client-side planar timewarp has to cover.
The server's pose-prediction horizon is the sum of its own pipeline and the
client-reported pipeline, but the client report guessed the display side as a
single refresh interval (8.3 ms at 120 Hz). The real decode-to-photon path —
the wait until the render loop picks the frame up, the in-flight buffer queue,
and the compositor's present pipeline — is typically 20-30 ms, so the horizon
under-predicted head position by that much and the residual landed on the
client warp as visible translation lag.

The renderer now measures actual decode-to-photon once per displayed frame
(drawable presentation time minus the frame's decode wall time, both in the
mach clock domain), and LatencyReporter carries an EMA of it in the report in
place of the one-refresh guess, with the guess kept as a fallback until the
first measurement. The previously always-zero displayedFrameAgeMs field is now
populated with the same measurement, feeding the ABR health signals and the
Home "Frame Age" stat.
SBudarin and others added 27 commits September 1, 2026 17:03
VideoToolbox invalidates a decompression session when the app loses its
foreground or immersive privilege. The decoder only counted and logged the
resulting kVTInvalidSessionErr, so every subsequent frame failed with the same
status forever while video kept arriving over the network — the headset sat on
a frozen frame until the server stopped sending. Rebuild the session from the
retained VPS/SPS/PPS instead, request a keyframe, and drop inter slices until
it lands, rate-limited so a dead session cannot trigger a rebuild per frame.

The client-side watchdogs could not see this at all: they only tracked NAL
units delivered by the receiver, which is a network counter. Video that
arrives but never decodes now trips its own five-second watchdog, which drops
the connection so discovery can re-establish it with a fresh decoder and a
fresh ARKit session, rather than the stall going unnoticed until the stream
ended and being reported as "Server stopped streaming".
A dropped connection sets both connectionState and wantsImmersiveSpace, which
notifies both of ContentView's observers, and each spawned its own
synchronizePresentationState task. Two overlapping dismissImmersiveSpace calls
could leave the space open with nothing left to drive it: the render loop then
ran forever against providers that disconnect had already stopped, querying a
device anchor it could never get and presenting drawables the compositor threw
away — pinning the GPU and filling the log with "this drawable won't be
presented". Serialize the transitions behind a single in-flight flag, with a
re-run if state changes while one is running, and route the disconnect button
through the same path so it dismisses before tearing the session down.

Harden the render loop against the same shape of failure: idle while nothing
is streaming, and skip submission for frames ARKit has no device anchor for
(the first frames after connecting, while the session is still coming up).
prepareForNewSession recreates the ARKit session and providers and cancels the
run task, but cancellation does not stop a runSession already suspended inside
makeProviders — requesting authorization and constructing Accessory(device:)
both take long enough for a disconnect/reconnect to swap the session and
providers underneath it. The stale task then resumed and ran a provider array
that no longer matched the live session, mixing an already-stopped provider
into it. In the field this appeared as hand tracking logging "provider is not
running" thousands of times across a whole session while world and accessory
tracking were fine, and only on some reconnects — a race.

Give each start a generation, bumped on every teardown, and re-check it after
every suspension point before running the session, storing the accessory
provider, or starting the anchor-consume task.
Restoring the SpatialGamepad controller profile took the Info.plist from a
state that predated NSLocalNetworkUsageDescription, dropping the key along
with it. Discovery still works on the current OS, but the string is what the
system shows when it asks for local network access, and nothing else supplies
it — it is not among the INFOPLIST_KEY_* entries in OXRSysVersion.xcconfig.
 visionOS: PSVR2 Sense controller input, seamless connect/reconnect, and stream recovery
The Viewer Settings sheet clipped its row labels on the left and its
values on the right: the default form style puts labels in a leading
column that does not fit the sheet's 420x360 minimum size.

Use the grouped form style on macOS, open the sheet at a size that fits
its content, and let the sliders span their rows. iOS is unchanged.
Fix the macOS simulator viewer settings layout
Only CHANGES.md conflicted: both sides appended a bullet to the 1.3.0
"Fixed" list. Kept both entries — the visionOS compositor-stall and
Blender Metal swapchain fixes from this branch, and the macOS simulator
"Viewer Settings" layout fix from upstream (PR #29).
Build & Test has been red since 31a9062 (2026-08-14); the last green run was
2026-07-18. Three problems, none caused by a code change.

ripgrep is not installed on the macOS runners. The runtime jobs call rg to
confirm a removed CMake option still fails configuration, so the step aborts
with "rg: command not found" before anything is built. The Apple clients job has
the opposite symptom from the same cause: its platform UI boundary lint reads
`if rg ...; then fail; fi`, so a missing rg makes the condition false and the job
passes without ever running the check. Both jobs now install it, rather than
rewriting the lints to grep, which would not be behaviour-preserving: rg respects
.gitignore and skips binaries, so the boundary lint would start matching build
output and vendored sources.

android-actions/setup-android installs a default package set including the
obsolete `tools` package, which no longer resolves and fails sdkmanager.

With the lint fixed the x86_64 job reaches its tests for the first time, and the
VideoToolbox encoder suite fails there. That runner reports no hardware encoder
at all, so VideoToolbox falls back to a software encoder that holds output until
the session is completed, and the suite waits for a callback that never comes.

VideoEncoder::FlushPendingFrames drains it. Frames reach VideoToolbox from a
Metal command-buffer completion handler rather than synchronously, so a single
VTCompressionSessionCompleteFrames can flush a pipeline the frame has not
reached yet; it retries until no frame is in flight. The session stays usable
and the callback drain keeps running, unlike Shutdown, which stops the drain
before it flushes and so discards the callbacks it produces. The test flushes
before waiting, and the suite now passes on both architectures - 2.14s on the
Intel runner that previously timed out.
The x86_64 Home test step intermittently fails with "the test runner hung
before establishing connection". The hang is not caused by the encoder
suite: it reproduced on a tree whose only delta from a green run was
comments, and again with ctest ordered after the Xcode steps so the suite
never ran. On failure, upload the xcresult (which records the hung
process diagnostic), crash/spin reports, and the syspolicyd/testmanagerd
log window to identify what blocks app launch on the Intel runners.
The x86_64 Home test step intermittently failed with "the test runner hung
before establishing connection". It was not the encoder suite: the hang
reproduced on a tree whose only delta from a green run was comments, and
again with ctest ordered after the Xcode steps so the suite never ran.

The spindump testmanagerd takes at the attach timeout shows the app's main
thread blocked for the whole 300s window constructing the launch-time
"Runtime is not configured" NSAlert: decoding its nib fetches an icon
through a synchronous IconServices XPC call, and iconservicesagent on the
Intel runner image is crash looping (SIGABRT in Metal's MTLLoader while
loading a RenderBox shader archive; the VM has no GPU), with launchd
throttling its respawn. The reply only arrives if the agent is allowed to
restart in time: the one green Intel run spent 240s of its 300s budget in
this same call, while arm64 runners have a working agent and attach in 9s.

The guidance is interactive first-launch UI for a person. Skip presenting
it when the app is hosting an XCTest session, so launch never depends on
IconServices and the test runner attaches immediately. Test behaviour is
unchanged: no test exercises the auto-presentation.
With the Home launch hang fixed, Runtime (x86_64) is the critical path at
534s: cmake build 175s, configure 88s, Home build+test 165s, simulator
build 49s.

ccache backs both Runtime jobs, covering C, C++, ObjC and ObjC++ so the
VideoToolbox encoder and swapchain sources are included. Keyed per commit with a
per-arch prefix fallback, and capped at 500M so it cannot evict the third-party
source cache under GitHub's 10GB LRU budget. A stats step reports the hit rate.
Warm, the arm64 cmake build drops from about two minutes to 8s.

The shared Swift packages cache their .build directories and the SwiftPM cache.
The key includes xcodebuild -version, because Swift module caches are
toolchain-specific: a runner image bump has to miss rather than restore modules
the new compiler rejects.

xcodebuild also gains COMPILER_INDEX_STORE_ENABLE=NO, which skips
index-while-building. Xcode Cloud sets this for CI already and nothing here
consumes the index store; it took the Home step from 165s to 108s.

-showBuildTimingSummary is added to measure where the Xcode time actually goes.
Those steps dominate the critical path once ccache is warm, and whether a
compilation cache would help depends on how much of them is compilation rather
than linking, asset catalogs and copy phases. The flag is for that measurement
and can come out once the answer is known.
With ccache warm, the two macOS xcodebuild steps dominate Runtime (x86_64):
Home build+test 110-130s, simulator build 53s, and -showBuildTimingSummary
shows Swift compilation is most of their task time (simulator target:
SwiftCompile 34.6s of ~55s).

COMPILATION_CACHE_ENABLE_CACHING keys compile jobs by content, not mtimes, so
unlike DerivedData caching it survives actions/checkout rewriting every mtime.
Each step gets its own CAS directory and cache entry: compilation caching is
reported unreliable when one cache serves several target contexts, and
per-step scoping also means one entry can be dropped without touching the
rest. The Apple clients job keeps its six destination builds uncached for the
same reason; it is off the critical path at under three minutes.

Verified locally: against a warm CAS, a build from empty DerivedData replays
136/136 compile jobs in 6.4s where the cold build took minutes. The CAS is
250M per target and arch; diagnostic remarks are on so hit rates are visible
in the job log.
The package job rebuilt everything from scratch each run, and since it is
serialized behind the other jobs its 2.5-3.5 minutes is pure tail on the
workflow's wall clock.

The third-party sources get the same _deps cache the Runtime jobs use, keyed
for the universal build directory. The Release Home build gets its own
compilation CAS; the build script owns the xcodebuild invocation, so the cache
settings arrive through XCODE_XCCONFIG_FILE rather than script changes. Xcode
compiles each arch separately even for ARCHS="arm64 x86_64", so both slices
replay from the CAS.

ccache is deliberately absent: the universal cmake build compiles with two
-arch flags per invocation, which ccache does not cache, so it would install
and run for nothing. The cmake compile remains the job's main cost.
…er helper

When the runtime dylib is dlopen'd in-process by an x86_64/Rosetta host (for
example CrossOver's Wine host), VideoToolbox refuses it the hardware HEVC
encoder: VTCompressionSessionCreate with RequireHardware=YES fails
kVTCouldNotFindVideoEncoderErr (-12908), and with Require=NO it silently
returns a software session (~27-40ms/frame, which saturates the CPU at full
resolution). Hardware H.264 is granted under Rosetta, but Quest clients only
decode HEVC. The only way to reach the hardware HEVC encoder is a native-arm64
process, so this adds one, behind an opt-in config key.

All GPU composition (blit/downscale/foveation) stays in the runtime; only
VTCompressionSessionEncodeFrame moves out of process. The runtime shares its
existing compose IOSurfaces with the helper zero-copy (mach send rights via
IOSurfaceCreateMachPort/IOSurfaceLookupFromMachPort) over a bootstrap
rendezvous performed once at startup, so no frame pixels cross the process
boundary. Per-frame "encode slot N" requests and the returned Annex-B NAL
units plus metrics ride an inherited Unix stream socket with explicit
little-endian framing; the IPC header is framework-free so it compiles
unchanged into both the x86_64 dylib and the arm64 helper. Timestamps cross as
int64 ns the parent supplies and the child echoes, because mach_absolute_time
is not comparable across the Rosetta boundary.

Fails safe. The helper is off by default (encoder_helper = false) and is
skipped for H.264 and for HEVC Main10, which it does not implement. If it
cannot start, does not get the hardware encoder, or dies mid-session, the
runtime logs where it broke and keeps using the in-process session. Frames in
flight to a stopped or dead helper are finalized as dropped so their slots and
callback-drain leases are returned, and the client object is released only
after VideoEncoder's callback drain is empty, so a Metal completion handler
racing teardown cannot use a freed client.

New:
  runtime/encoder_helper/EncoderHelperIpc.h   framework-free wire protocol
  runtime/encoder_helper/main.mm              arm64 helper (VT hardware HEVC)
  runtime/encoder_helper/CMakeLists.txt       standalone arm64 target
  runtime/encoder_helper/build-helper.sh      single-file arm64 build
  runtime/encoder_helper/smoke_test.mm        offline cross-arch end-to-end test
  runtime/encoder_helper/README.md            design, build, enable, test
  runtime/src/HevcEncoderHelperClient.{h,mm}  spawn / IPC / surface glue
  tests/TestEncoderHelperIpc.cpp              wire-format coverage
Changed:
  runtime/src/VideoEncoder.{h,mm}  delegate the VT encode when the helper is up
  runtime/src/Config.{h,cpp}       encoder_helper / encoder_helper_path keys
  runtime/oxrsys-runtime.toml      document the new keys
  runtime/CMakeLists.txt           build the client, link IOSurface

Verified on an M4 Pro: the bundled smoke test drives an x86_64 (Rosetta) parent
against the arm64 helper and reports InitAck hardware=YES with an isolated
hardware HEVC encode of avg 8.36ms/frame at 2272x1264. Runtime dylib builds
x86_64 (lipo-verified), helper builds arm64 via both build-helper.sh and its
CMakeLists, and the test suite passes 586 assertions in 65 cases.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018yQ8zUuprBfVcNWabKkVVM
…asured hardware availability

The helper added in the parent commit was gated to HEVC Main 8-bit and enabled
only by hand. It now covers every codec the runtime negotiates, and the runtime
picks the path on its own.

Codec coverage. Init carries the negotiated codec and profile (protocol v2), and
the helper builds its VideoToolbox session from them: H.265 Main, H.265 Main10
and H.264 Main, with the profile requested exactly as the in-process path
requests it and the same log-and-continue fallback when a profile is
unavailable. Parameter-set emission picks the H.264 or HEVC CoreMedia accessor;
the Annex-B framing, the BGRA surface contract and the IPC are unchanged.
Main10 needs no format work: develop already encodes Main10 from the 8-bit BGRA
compose surface (a 10-bit bitstream from an 8-bit source), so the helper does
the same. True 10-bit source would mean a new compose format end to end and is
deliberately not attempted. Both peers now reject a frame whose protocol version
is not theirs, so a stale helper binary fails the handshake instead of
misreading a payload.

Automatic selection. runtime/src/EncoderPathPolicy.{h,cpp} decides once per
session from (negotiated codec, profile, whether THIS process can get a hardware
encoder for it, the encoder_helper override). Availability is asked of
VideoToolbox - VTCopyVideoEncoderList filtered by CodecType +
IsHardwareAccelerated, falling back to a RequireHardware=YES session probe - not
inferred from the process architecture, so the logic survives Apple changing
what Rosetta is granted. The helper is used only where the in-process encoder
would be software: a natively running runtime, and a Rosetta host negotiating
H.264 (which Rosetta is granted), both keep the in-process path rather than
paying for a process hop. AV1 never routes to the helper - honouring the
negotiated codec outranks reaching hardware.

encoder_helper becomes tri-state ("auto" default, "true"/"false" to force a path
for debugging) and the fail-safe behaviour is unchanged: any helper failure logs
where it broke and falls back to the in-process session.

RequireHardwareAcceleratedVideoEncoder is now YES on the in-process path when
the query says this process can have it, with a warn-and-retry without the
requirement if the create still fails, so a session is never lost over it. The
runtime also logs whether the session it ended up with is hardware or software.
The unconditional NO is what made the Rosetta software fallback silent.

Measured on an M4 Pro, 2272x1264, 35 frames after warmup, both paths from the
same IOSurfaces with matched session properties and thread QoS:

  parent            codec         in-process          helper (round trip)
  x86_64 (Rosetta)  H.265 Main    19.5ms software     8.3ms hardware
  x86_64 (Rosetta)  H.265 Main10  33.6ms software     8.2ms hardware
  x86_64 (Rosetta)  H.264 Main     7.8ms hardware     8.0ms hardware
  arm64 (native)    H.265 Main     8.4ms hardware     8.3ms hardware
  arm64 (native)    H.265 Main10   8.4ms hardware     8.6ms hardware
  arm64 (native)    H.264 Main     8.2ms hardware     7.9ms hardware

The IPC round trip costs 0.03-0.06ms, so the boundary is not the cost; the
helper is worth it exactly where the in-process encoder is software, which is
what the policy encodes.

smoke_test.mm now encodes both paths from the same surfaces and reports each,
with --codec / --profile / --frames and per-path skips; build it arm64 as well
as x86_64 for the native-host numbers. HevcEncoderHelperClient is renamed
EncoderHelperClient, since it is no longer HEVC-only.

Tests: 720 assertions in 76 cases (was 586/65). TestEncoderPathPolicy covers the
selection policy as a pure function - no GPU needed - including that Auto never
routes to the helper when in-process hardware exists and that AV1 never routes
at all; TestEncoderHelperIpc covers the v2 Init payload and version/bounds
rejection; TestVideoToolboxEncoder gains Main10 coverage. The VideoToolbox
encoder tests were also run x86_64 against a real arm64 helper (automatic
selection, forced selection, and the missing-binary fallback) and arm64 with the
helper forced on for H.264 and Main10. Runtime dylib builds x86_64
(lipo-verified); helper builds arm64 via build-helper.sh and its CMakeLists.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018yQ8zUuprBfVcNWabKkVVM
… software encoder

oxrsys_videotoolbox_encoder_tests failed 3 of 4 cases on GitHub's
macos-26-intel runner. The runner advertises no hardware video encoder, so
H.264 falls back to a software VideoToolbox session, and a software session
buffers a single submitted frame until the compression session is flushed -
it will never volunteer it. EncodeOneFrame submitted exactly one frame and
then blocked on a passive REQUIRE(wait_for(5s)), so it timed out every time.

The tell was in the failing log itself: the completion callback fired at
+5.05s, immediately after the timeout, because the REQUIRE threw past the
explicit Shutdown and the encoder's destructor flushed the session on the way
out. That is also why "VideoToolbox shutdown drains submitted frame-source
ownership" was the one case of four that passed - same runner, same codec,
same encoder, but it calls Shutdown (VTCompressionSessionCompleteFrames)
rather than waiting. The "mutex lock failed: Invalid argument" warnings were
that late callback landing on a mutex Catch2 had already destroyed, i.e.
teardown fallout rather than the cause.

EncodeOneFrame now drains the encoder before asserting: it keeps the wait, so
a backend that answers on its own is unaffected, and only falls back to an
explicit Shutdown if the wait expires. The lock is released across that call
because both callbacks take it from VideoToolbox's own thread, and Shutdown
returns only once every callback has run. A second Shutdown short-circuits on
resourcesDestroyed_, so the existing Shutdown below it stays correct. The
callback state is also hoisted above the encoder so it outlives it, which
removes the destroyed-mutex teardown race outright.

The timeout is no longer a REQUIRE: on a software backend it is an expected
path, not a failure. What is asserted is that the frame encoded - CHECK on
completed, not-dropped, NAL count and Annex-B framing - not how promptly a
particular backend chose to hand it back. The grace before the flush is 750ms
rather than 5s: hardware sessions and the out-of-process helper both answer in
single-digit milliseconds, and now that a timeout no longer aborts the case,
every iteration of the format loop runs, so a 5s grace would have pushed the
suite past its own 30s ctest timeout.

This is pre-existing on develop, which carries the identical wait. It was
never observed because the Runtime lanes died at step 1 on `rg: command not
found` and had never once reached the Test step.

Verified by forcing a genuinely software session locally (a throwaway
EnableHardwareAcceleratedVideoEncoder=@no build): without this change the
local run reproduces the runner exactly, 4 cases / 1 passed / 3 failed and
31 assertions / 3 failed; with it, 90 assertions in 4 test cases pass on both
arm64 and x86_64. Also re-run against a real arm64 encoder helper with
encoder_helper forced on, covering H.264 and H.265 Main10 through the
out-of-process path, where every frame lands well inside the grace and the
flush fallback is never entered.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018yQ8zUuprBfVcNWabKkVVM
Speed up CI with compiler and Swift package caching
…time

The helper only built through its own build-helper.sh / standalone CMake
project, so nothing in the standard flow produced it: `cmake --build`, the
ctest lane and scripts/macos_build_package.sh all shipped a runtime whose
helper lookup (a sibling of the dylib) found nothing and silently fell back.

runtime/CMakeLists.txt now adds runtime/encoder_helper, and the
oxrsys-encoder-helper target:

* is always arm64, via its own OSX_ARCHITECTURES, whatever the runtime is
  configured for (x86_64, arm64 or universal) - escaping Rosetta is the
  point of it, and it links only system frameworks so no x86_64
  FetchContent artifact can reach its link;
* is written next to liboxrsys-runtime.dylib, where VideoEncoder looks for
  it when encoder_helper_path is empty;
* is a dependency of oxrsys_runtime, so the packaging scripts' `--target
  oxrsys_runtime` build produces it too;
* is ad-hoc signed after linking.

The standalone configure (cmake -S runtime/encoder_helper) and
build-helper.sh keep working for iterating on the helper alone.

ctest runs oxrsys_videotoolbox_encoder_tests with OXRSYS_ENCODER_HELPER_PATH
pointing at the build's own helper (and a 60 s timeout for the helper
lifecycle cases that follow), instead of the test binary's own directory.

macos_build_package.sh copies the helper into runtime/ beside the dylib and
fails on a helper that is not arm64-only; macos_sign_notarize.sh signs it
with the hardened runtime (a separately spawned executable needs its own
Developer ID signature to notarize) and puts it in the archive. Docs list
the new package entry. No workflow change is needed: the Runtime job's
`cmake --build` and the package job's script both pick it up.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018yQ8zUuprBfVcNWabKkVVM
… BT.709 in the helper

Three review findings on the out-of-process encoder helper, each with a
regression test that fails (or crashes) without the fix.

1. SIGPIPE on the control socket. A write to a helper that had died raised
   SIGPIPE, whose default action kills the writer - the game process, since
   the runtime dylib lives inside it under Wine. Reproduced: with the
   client's reader parked in a callback and the helper SIGKILLed, the next
   SubmitFrame killed the test binary with signal 13. The helper had the
   mirror problem: the host closing on frames in flight killed it with
   SIGPIPE while it flushed them (wait status 13).

   Fixed per socket, never with a process-wide SIG_IGN in the host:
   SO_NOSIGPIPE on both ends of the socketpair plus MSG_NOSIGNAL on every
   send, in one shared enc_ipc::SendAll/RecvAll/DisableSigPipe used by the
   client, the helper and smoke_test. A dead peer is now EPIPE, which marks
   the helper dead and falls back in-process. The helper, being its own
   process, also ignores SIGPIPE, which covers its stderr pipe to a host
   that is gone.

   While there: MarkDead closed the socket from whichever thread noticed the
   death while the reader thread could still be in read() on it, so the
   reader could land on a descriptor number the host had reused. It now
   shuts the socket down (which still wakes the reader and EOFs the child)
   and the fd is closed only in Stop(), after the reader is joined.

2. Use-after-free of context->slotIndex. EncodeInternal's completion handler
   published the frame context into helperContexts_ and only then read
   context->slotIndex for the submit. Once published, the helper reader
   thread may free it at any time: OnHelperDied -> ReclaimHelperFrames
   finalizes every published context. With the window widened by a sleep,
   ASan reports heap-use-after-free at exactly that read on every run; with
   the fix and the same sleep it is clean. The slot is now captured before
   publishing and `context` is never dereferenced after it.

   Related hardening: the cookie was the context's address, which the
   allocator can hand to the next frame, so the "orphan" lookup after a
   failed submit could in principle finalize a different frame. Cookies are
   now a per-encoder sequence number. SubmitFrame now reports whether the
   request was written: only a frame that was never sent is reclaimed by the
   submitter; one that was sent is left to FrameDone or OnHelperDied, so the
   two paths can no longer race over a frame the reader is still delivering.

3. BT.709 in the helper. The in-process session sets ColorPrimaries,
   TransferFunction and YCbCrMatrix to ITU_R_709_2; the helper's session set
   none, and its SPS carried no VUI at all (checked with an independent SPS
   parser on H.264 Main, HEVC Main and HEVC Main10). The contract now lives
   once in encoder_helper/EncoderSessionColor.h, applied by the in-process
   session, the helper and smoke_test, so the paths cannot drift again.
   Helper output now carries video_format 5, primaries/transfer/matrix 1/1/1,
   video range - identical to the in-process hardware session.

Tests:
* TestEncoderHelperIpc: a vanished or shut-down peer is EPIPE, not SIGPIPE
  (asserts SIGPIPE is at SIG_DFL, so a regression kills the binary).
* TestEncoderHelperClient (real helper): host survives the helper dying with
  a write pending and gets exactly one died callback; the helper exits
  normally when the host hangs up on frames in flight; the helper's SPS VUI
  (parsed by CoreMedia from the bitstream) is BT.709 video range for H.264
  Main, HEVC Main and HEVC Main10. Each SKIPs when the helper cannot come up
  on a hardware encoder (Intel, or a runner without one).
* TestVideoToolboxEncoder: every encoded stream, whichever path produced
  it, carries the BT.709 colour description; and when the helper is live,
  killing it mid-stream reclaims the in-flight frames, keeps encoding
  in-process, finalizes no frame twice and still shuts down cleanly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018yQ8zUuprBfVcNWabKkVVM
VideoToolbox's software HEVC encoder writes full-range video for a BGRA
source once HEVC Main is requested explicitly, while every hardware
encoder writes video range. The software session is what an x86_64
runtime under Rosetta falls back to when the encoder helper dies, so a
helper death flipped the stream's range mid-session, against the
limited-range stream contract. VideoToolbox has no compression property
for range; a session takes it from its source.

A software in-process session is now never handed the BGRA compose
surface. Each frame is first converted by a VTPixelTransferSession into
an explicitly video-range 420YpCbCr8BiPlanarVideoRange buffer with the
same BT.709 matrix, primaries and transfer function, and the encoder
keeps that range. The conversion is set up in Initialize, not on helper
death, and costs about 0.3 ms per 2272x1264 frame under Rosetta against
20-34 ms for the software encode. Hardware sessions and the helper are
unchanged. The contract lives in EncoderSessionColor.h with the rest of
the colour properties.

Tests: a test-only switch pins the in-process session to the software
encoder on any machine. The codec test now asserts video range on every
path, the new software case covers H.264, HEVC Main and HEVC Main10, and
the helper-death case checks the range of the frames encoded after the
kill. With the conversion disabled, the software HEVC Main case fails on
both architectures and the helper-death case fails on x86_64.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018yQ8zUuprBfVcNWabKkVVM
Record the encoder helper where AGENTS.md asks significant changes to be
recorded: contract bullets in AGENTS.md for the encode-path policy, the
IPC and Mach-port boundary, the helper-death fallback and SIGPIPE-safe
sockets, the shared colour contract, and arm64-only packaging; release
notes in CHANGES.md under the pending 1.3.0 release; the encode-path
decision and fallback in docs/architecture.md; the encoder_helper and
encoder_helper_path keys with their defaults in the macOS Home config
reference; the test coverage and the physical-headset gate in
docs/testing-and-conformance.md; a support-matrix row; and one clause
on the existing README streaming highlight.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018yQ8zUuprBfVcNWabKkVVM
…lper

Hardware encoder helper for all codecs, with the encode path chosen automatically from measured availability
@demonixis
demonixis merged commit f6c4044 into main Oct 3, 2026
5 checks passed
@demonixis
demonixis deleted the develop branch October 3, 2026 20:13
@demonixis
demonixis restored the develop branch October 3, 2026 20:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants