Repository navigation
Conversation
nburns
marked this pull request as draft
September 22, 2026 23:54
The foveated encode placed its high-resolution region at a fixed centre, so eye tracking on the client had nothing to drive. Clients may now report an eye gaze direction and the server steers the foveal centre to follow it. Gaze arrives as TrackingPacket.gazeDirection, a head-space unit vector gated by TRACKING_FLAG_EYE_GAZE_ACTIVE and appended at the end of the struct so clients predating it keep working. The server maps it in tan space against the client's reported eye FOV and low-passes it, since following fixation jitter directly would shimmer the foveal boundary every frame. The client has to un-warp with exactly the centre the server warped with, or it reconstructs a geometrically wrong image rather than a merely stale one. So the server quantizes the requested shift first and warps with the dequantized result, making the transmitted bytes precisely what produced the image. Both ends then call the same AlignCenterShift. The bytes reuse previously reserved space in VideoPacketHeader, TcpVideoNalHeader and TcpRenderPose under VIDEO_FLAG_FOVEATION_CENTER, so no wire size changes except TrackingPacket. Moving the centre never changes the encoded resolution, because the optimized size depends on centre size and edge ratio but not centre shift. It therefore costs no encoder reconfigure and is safe to update every frame. Clients that report no gaze keep the previous fixed-centre behaviour exactly. Two fixes fell out of this: AlignCenterShift rounded away from zero, so a requested shift of exactly 1.0 aligned to 1.0077 and pushed the warp outside [0, 1], breaking the endpoint invariant the shader relies on. Harmless while the shift was always zero. The TCP tracking path required a full-size TrackingPacket, duplicating a check that TrackingReceiver::InjectPacket already performs correctly against kMinTrackingPacketSize. Growing the struct would have silently dropped all tracking from older clients over TCP.
Review fixes for gaze-driven foveated encoding, plus tests that pin the server/client contract. Correctness and robustness: - Steer only for clients that advertise the new CLIENT_CAPABILITY_FOVEATION_CENTER. A client that cannot un-warp a per-frame centre must keep the static centre it was announced, or it reconstructs a geometrically wrong image; previously any gaze-reporting client was steered. Gating also skips the per-frame tracking copy and tan math entirely for non-participating clients. - Reject non-finite gaze vectors and eye FOVs before use, and make QuantizeCenterShift refuse NaN. A single poisoned packet previously parked the smoothing accumulator on NaN for the life of the process. - Reset the smoothed centre on connect, disconnect, and server start, and decay it to zero shift while gaze is inactive instead of freezing at the last fixation. The old state silently carried one session's fixation into the next client's stream. - Fix the gaze-to-shift mapping. The warp applies one un-mirrored shift to both eye halves, so the target is the average of the two mirrored eyes' fixation points: the horizontal FOV asymmetry cancels and the vertical axis keeps its recentring term. The shader places the sharp region's centre at UV 0.5 + shift * (1 - centerSize) / 2, so the mapping needs the 1 / (1 - centerSize) scale; without it the fovea travelled only ~40-55% of the way to the fixation point on the shipped presets. - Round AlignCenterShift to nearest instead of ceil, which biased every off-grid value toward +1 by up to one grid cell. The old comment's endpoint-overshoot story was wrong: the clamp guards float error, and +-1.0 land within an ulp of exact. - Keep the static preset shift until a gaze centre has actually been set (tracked by a set bit in the packed atomic), instead of silently discarding foveationSettings_.centerShiftX/Y, and only stamp VIDEO_FLAG_FOVEATION_CENTER when steering really produced a centre. - Share one DecodeCenterShift between the encoder's warp and the client-side ApplyQuantizedCenterShift so the reconstruction is bit-identical by construction, and stamp all four video header builders (render pose, TCP NAL, UDP NAL, FEC parity) through a single StampFoveationCenter helper; the parity header previously omitted the centre. Tests: - TestGazeFoveation.cpp: a line-for-line C++ replica of the Metal kernel's compress_axis proves the steered foveal centre lands on the gaze fixation point for asymmetric FOVs on both axes, through the full quantize -> wire -> decode pipeline. Also: a bit-exact server-warp vs client-un-warp round trip over every possible wire byte and preset, alignment and quantization properties, rejection of inactive/degenerate/non-finite input, and filter convergence/decay/reset. - Pin the repurposed reserved bytes with offsetof asserts in TestProtocolLayout.cpp (TcpVideoNalHeader, TcpRenderPose), matching the existing VideoPacketHeader pins. - Bring the Swift protocol mirror back in sync: foveationCenterX/Y, hasFoveationCenter, VideoFlags.foveationCenter, the capability bit, TrackingPacket.gazeDirection, and layout assertions for all of them (size 1076 / stride 1080). The vertical sign convention still needs eye-tracking hardware to verify; the mapping and its rationale now live in runtime/src/GazeFoveation.h where both the smoothing filter and the offset computation are unit-tested.
nburns
force-pushed
the
gaze-foveation
branch
from
September 23, 2026 16:22
ad3b42b to
a04e321
Compare
Contributor
Author
|
One design point worth explicit sign-off rather than sliding through review: Alternatives if you'd rather keep the bit free:
The flag felt most conventional, which is why the PR does that — but it is the last one, so flagging the trade-off for a deliberate call. Happy to switch to either. |
…d frame Un-warping with the most recently received centre instead of the displayed frame's own centre reconstructs a geometrically wrong image whenever the centre is moving: decode pipelines run several frames deep, so the latest value leads the frame on screen and the periphery visibly stretches. State the client obligation explicitly - match by frame index or presentation time - since every value in that failure is individually correct and the bug is invisible while the centre holds still.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Lets clients report an eye-gaze direction and steers the foveated-encode centre
to follow it, instead of the fixed centre the presets announce.
How it works
TrackingPacket.gazeDirection(head space, gated byTRACKING_FLAG_EYE_GAZE_ACTIVE), appended to the struct so older clients'shorter packets still parse. Non-finite values are rejected.
CLIENT_CAPABILITY_FOVEATION_CENTER: a clientthat cannot un-warp a per-frame centre must keep the announced static centre,
or it reconstructs a geometrically wrong image. Without the bit, nothing
changes.
runtime/src/GazeFoveation.h) works in tan space against theclient's eye FOV, targets the average of the two mirrored eyes' fixation
points (one un-mirrored shift serves both halves), and scales by
1 / (1 - centerSize)to match where the warp kernel actually places thesharp region. The result is low-passed; on gaze dropout it decays back to
centre, and it resets across sessions.
int8, warps with the dequantized value,and transmits those exact bytes. Both ends call the same
DecodeCenterShift,so client reconstruction is bit-identical by construction.
Wire impact
Only
TrackingPacketgrows (appended field). The centre reuses reserved bytesin
VideoPacketHeader,TcpVideoNalHeader, andTcpRenderPose, flagged byVIDEO_FLAG_FOVEATION_CENTERand stamped by one shared helper in every headerbuilder. Moving the centre never changes the encoded size, so it costs no
encoder reconfigure and can update every frame.
Also fixed along the way:
AlignCenterShiftrounds to nearest (ceil biased thecentre toward +1 by up to a grid cell), and the TCP tracking path no longer
duplicates the size check
TrackingReceiver::InjectPacketalready does, whichwould have dropped older clients' tracking once the struct grew.
Testing
oxrsys_runtime_testspasses.TestGazeFoveation.cppchecks the steeringgeometry against a C++ replica of the Metal kernel (asymmetric FOVs, both axes,
full quantize -> wire -> decode pipeline) and round-trips every possible wire
byte per preset, asserting server warp == client un-warp bit-exactly. Layout
tests on both the C++ and Swift side pin every repurposed byte; the Swift
protocol mirror is updated in lockstep.
Verified on macOS arm64 (Metal/VideoToolbox). Not yet verified against a real
eye: the vertical sign convention is self-consistent in code and tests, but
only eye-tracking hardware can confirm the sharp region tracks vertical gaze
rather than mirroring it. Happy to flip it if a maintainer with hardware sees
otherwise.