Skip to content

Add a Steam Frame client (draft) - #30

Draft
nburns wants to merge 10 commits into
demonixis:developfrom
nburns:steam-frame-client
Draft

nburns wants to merge 10 commits into
demonixis:developfrom
nburns:steam-frame-client

Conversation

@nburns

@nburns nburns commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

A from-scratch Vulkan + OpenXR streaming client for Valve's Steam Frame, paired with the OXRSys runtime (clients/steam-frame/). Reuses the in-tree oxrsys/protocol lib and transport. Draft — everything verifiable on macOS is done; on-device bring-up needs the hardware.

Done (verified on macOS vs. a live OXRSys server)

  • Stereo OpenXR + Vulkan (two per-eye swapchains → projection layer), 6DoF head pose
  • Extension-enumeration dump (on-device graphics-binding confirmation)
  • FFmpeg HEVC/H.264 → Vulkan texture, per-eye SBS sampling
  • Live receive (discovery → ClientConnect → NAL feed) over the ported transport
  • Motion-to-photon loop: pose upstream + render-pose async timewarp + stale-frame reproject
  • Controllers (KHR-simple + Oculus-Touch), XR_EXT_hand_tracking, XR_EXT_eye_gaze_interaction
  • Foveated-encoding (AADT) decode — un-warps FFE video (~21% smaller frame, read-back verified)
  • Adaptive-bitrate LatencyReport; keyframe request on decode error
  • Headless tests (ctest): decoder output + XOR-FEC recovery

Remaining (needs the Frame)

  • Build Linux ARM64 on the Sniper runtime; deploy via SteamOS Devkit Client
  • Switch graphics binding to XR_KHR_vulkan_enable2
  • Hardware decode (Vulkan Video / VAAPI) — replaces ~110 ms/frame software decode
  • Verify real gaze / hand / controller data + loader active_runtime path on-device
  • Reprojection comfort verdict under real motion + network jitter

Follow-up

  • Factor the shared transport (NetworkReceiver / TrackingSender, de-Androidized from clients/android-vr) into clients/shared/ so both clients use one copy

Full build/run notes: clients/steam-frame/README.md.

@demonixis

Copy link
Copy Markdown
Owner

Thanks for this contribution!

@nburns nburns mentioned this pull request Sep 21, 2026
@demonixis

Copy link
Copy Markdown
Owner

Hey, I plan to push a release soon (in the next coming days). Do you want that we wait for your PR?

@nburns

nburns commented Sep 22, 2026

Copy link
Copy Markdown
Contributor Author

Still no hardware, so I certainly wouldn't wait for this PR

🤨 Screenshot 2026-09-21 at 1 03 26 PM

The foveated encode placed its high-resolution region at a fixed centre, so
eye tracking on the client had nothing to drive. Clients may now report an eye
gaze direction and the server steers the foveal centre to follow it.

Gaze arrives as TrackingPacket.gazeDirection, a head-space unit vector gated by
TRACKING_FLAG_EYE_GAZE_ACTIVE and appended at the end of the struct so clients
predating it keep working. The server maps it in tan space against the client's
reported eye FOV and low-passes it, since following fixation jitter directly
would shimmer the foveal boundary every frame.

The client has to un-warp with exactly the centre the server warped with, or it
reconstructs a geometrically wrong image rather than a merely stale one. So the
server quantizes the requested shift first and warps with the dequantized
result, making the transmitted bytes precisely what produced the image. Both
ends then call the same AlignCenterShift. The bytes reuse previously reserved
space in VideoPacketHeader, TcpVideoNalHeader and TcpRenderPose under
VIDEO_FLAG_FOVEATION_CENTER, so no wire size changes except TrackingPacket.

Moving the centre never changes the encoded resolution, because the optimized
size depends on centre size and edge ratio but not centre shift. It therefore
costs no encoder reconfigure and is safe to update every frame. Clients that
report no gaze keep the previous fixed-centre behaviour exactly.

Two fixes fell out of this:

AlignCenterShift rounded away from zero, so a requested shift of exactly 1.0
aligned to 1.0077 and pushed the warp outside [0, 1], breaking the endpoint
invariant the shader relies on. Harmless while the shift was always zero.

The TCP tracking path required a full-size TrackingPacket, duplicating a check
that TrackingReceiver::InjectPacket already performs correctly against
kMinTrackingPacketSize. Growing the struct would have silently dropped all
tracking from older clients over TCP.
Review fixes for gaze-driven foveated encoding, plus tests that pin the
server/client contract.

Correctness and robustness:

- Steer only for clients that advertise the new
  CLIENT_CAPABILITY_FOVEATION_CENTER. A client that cannot un-warp a
  per-frame centre must keep the static centre it was announced, or it
  reconstructs a geometrically wrong image; previously any gaze-reporting
  client was steered. Gating also skips the per-frame tracking copy and
  tan math entirely for non-participating clients.
- Reject non-finite gaze vectors and eye FOVs before use, and make
  QuantizeCenterShift refuse NaN. A single poisoned packet previously
  parked the smoothing accumulator on NaN for the life of the process.
- Reset the smoothed centre on connect, disconnect, and server start,
  and decay it to zero shift while gaze is inactive instead of freezing
  at the last fixation. The old state silently carried one session's
  fixation into the next client's stream.
- Fix the gaze-to-shift mapping. The warp applies one un-mirrored shift
  to both eye halves, so the target is the average of the two mirrored
  eyes' fixation points: the horizontal FOV asymmetry cancels and the
  vertical axis keeps its recentring term. The shader places the sharp
  region's centre at UV 0.5 + shift * (1 - centerSize) / 2, so the
  mapping needs the 1 / (1 - centerSize) scale; without it the fovea
  travelled only ~40-55% of the way to the fixation point on the
  shipped presets.
- Round AlignCenterShift to nearest instead of ceil, which biased every
  off-grid value toward +1 by up to one grid cell. The old comment's
  endpoint-overshoot story was wrong: the clamp guards float error, and
  +-1.0 land within an ulp of exact.
- Keep the static preset shift until a gaze centre has actually been
  set (tracked by a set bit in the packed atomic), instead of silently
  discarding foveationSettings_.centerShiftX/Y, and only stamp
  VIDEO_FLAG_FOVEATION_CENTER when steering really produced a centre.
- Share one DecodeCenterShift between the encoder's warp and the
  client-side ApplyQuantizedCenterShift so the reconstruction is
  bit-identical by construction, and stamp all four video header
  builders (render pose, TCP NAL, UDP NAL, FEC parity) through a single
  StampFoveationCenter helper; the parity header previously omitted the
  centre.

Tests:

- TestGazeFoveation.cpp: a line-for-line C++ replica of the Metal
  kernel's compress_axis proves the steered foveal centre lands on the
  gaze fixation point for asymmetric FOVs on both axes, through the
  full quantize -> wire -> decode pipeline. Also: a bit-exact
  server-warp vs client-un-warp round trip over every possible wire
  byte and preset, alignment and quantization properties, rejection of
  inactive/degenerate/non-finite input, and filter
  convergence/decay/reset.
- Pin the repurposed reserved bytes with offsetof asserts in
  TestProtocolLayout.cpp (TcpVideoNalHeader, TcpRenderPose), matching
  the existing VideoPacketHeader pins.
- Bring the Swift protocol mirror back in sync: foveationCenterX/Y,
  hasFoveationCenter, VideoFlags.foveationCenter, the capability bit,
  TrackingPacket.gazeDirection, and layout assertions for all of them
  (size 1076 / stride 1080).

The vertical sign convention still needs eye-tracking hardware to
verify; the mapping and its rationale now live in
runtime/src/GazeFoveation.h where both the smoothing filter and the
offset computation are unit-tested.
A parity group covered a run of consecutive packets, so any two adjacent losses
fell in the same group and one XOR parity could not recover them. Adjacent loss
is the normal case on Wi-Fi, where loss arrives in bursts rather than uniformly,
which left FEC recovering little of what actually goes wrong.

Interleaving assigns group g the packets g, g+groupCount, g+2*groupCount, ... so
neighbouring packets belong to different groups. The same single parity and the
same overhead then recover any burst up to groupCount packets long: burst
tolerance becomes the group count instead of one, for no extra bandwidth.

Modelled at a 1% burst-start rate on a 250-packet frame, whole-frame recovery
goes from 0.108 to 0.752 for two-packet bursts and 0.081 to 0.649 for
three-packet bursts. Doubling parity instead, without interleaving, only reaches
0.135 on the three-packet case for twice the overhead, so the layout is what
matters here rather than the parity count.

The layout is negotiated through CLIENT_CAPABILITY_FEC_INTERLEAVED and
SERVER_FEATURE_FEC_INTERLEAVED rather than assumed from a version. Every client
implements the grouping independently, and a receiver using a different layout
than the sender XORs a packet out of the wrong group, handing plausible garbage
to the decoder instead of failing cleanly. Clients that do not advertise the
capability keep the contiguous layout exactly.

Interleaved parity is emitted after a frame's data packets, because an
interleaved group is not complete until the frame is.

tests/TestProtocolFec.cpp covers this at the protocol level rather than inside
any one client: both layouts partition the frame exactly once, no group exceeds
FEC_GROUP_SIZE (receivers gather groups into fixed-size arrays), interleaving
puts every pair of neighbours in different groups, and a burst the contiguous
layout loses is recovered up to the group count and no further.

The android-vr client opts in. The shared Swift package gains the mirrored
layout and a flag, but stays on the contiguous layout: turning it on needs a
visionOS build to verify, which this change cannot do.
Review fixes for the interleaved FEC layout, plus receiver-level tests.

Correctness and robustness:

- Attempt recovery whenever the parity held could cover everything
  missing, instead of only when a single packet is missing. The old gate
  deferred every burst recovery -- the case interleaving exists for --
  to the next frame's first packet, delivering the frame one interval
  late; that flush path also discarded the triggering packet, seeding
  the next frame with a self-inflicted loss. Fixed in the Android and
  Swift receivers alike.
- Match parity packets by totalPackets as well as frameIndex. One
  frame's NALs share a frameIndex and each numbers its packets and
  groups from zero, so a reordered parity packet from a small NAL
  (SPS/PPS) could occupy the next NAL's parity slot and XOR-recover
  garbage. A data packet with a different totalPackets under the same
  frameIndex likewise starts a new reassembly now.
- Keep the inline parity schedule for contiguous connections: a
  contiguous group is complete mid-frame, and emitting all parity as
  one tail burst both exposed it to a single burst-loss event and
  changed behaviour for clients that never opted into interleaving.
  Only interleaved connections pay the end-of-frame cost. Both paths
  share one parity-send helper using fixed arrays.
- Renumber CLIENT_CAPABILITY_FEC_INTERLEAVED to 0x00001000: 0x00000800
  is claimed by CLIENT_CAPABILITY_FOVEATION_CENTER in the gaze
  foveation PR, and the collision would negotiate each feature on for
  clients advertising the other.
- Latch fecInterleaved_ once at SendNalUnit entry alongside the client
  snapshot, hoist the negotiation store out of the codec retry loop in
  both connect handlers, and reset it on disconnect and server start.
- Make the Android receiver's flag atomic and route the Swift one
  through the state lock; both receivers latch the layout per frame so
  a mid-frame flip can never mix layouts within one recovery.
- Guard GroupLayout::GroupOf against zero-packet layouts (division by
  zero) in C++ and Swift, add the members == 0 skip the Swift recovery
  already had, and delete the dead contiguous-only fec::GroupRange.
- Move the Swift capability/feature bits into ClientCapabilityFlags and
  ServerFeatureFlags, where clients actually build ClientConnect from.

Tests:

- Extract the frame reassembly and FEC recovery state machine from
  NetworkReceiver into a shared, header-only VideoFrameAssembler under
  common/streaming, so the receiver half of the FEC contract exists and
  is tested exactly once (in oxrsys_runtime_tests, on macOS) instead of
  being forked per UDP client. TestVideoFrameAssembler.cpp covers:
  burst recovery on parity arrival with no next-frame packet needed,
  the unrecoverable-burst NACK path without losing the next frame's
  trigger packet, cross-NAL parity rejection, inline contiguous
  recovery, short-last-packet size restoration, same-frameIndex NAL
  flushing, and duplicate data/parity handling.
- Pin SERVER_FEATURE_FEC_INTERLEAVED and CLIENT_CAPABILITY_FEC_INTERLEAVED
  values in TestProtocolLayout.cpp and the Swift ProtocolLayoutTests,
  and mirror the C++ GroupLayout property tests (exactly-once
  partition, member cap, neighbour separation, empty-layout inertness)
  in the Swift package so the two implementations cannot drift.
- Document the emission schedules and the totalPackets matching rule in
  docs/protocol.md, and move the FEC section out of the middle of the
  video-flag bullet list.
A from-scratch Vulkan + OpenXR streaming client for Valve's Steam Frame, paired
with the OXRSys runtime. Hosted-class, pure portable C++; builds against a
Metal/Vulkan OpenXR runtime on macOS for development and retargets to the
Frame's runtime by swapping loader discovery.

Reuses the in-tree oxrsys/protocol lib and transport (NetworkReceiver /
TrackingSender, de-Androidized from clients/android-vr). Implements stereo
projection, FFmpeg decode, live receive, a head-pose motion-to-photon loop with
render-pose async timewarp, controllers + hand tracking + eye gaze, foveated-
encoding (AADT) decode, adaptive-bitrate latency reports, and keyframe loss
recovery. Headless tests cover the decoder and XOR-FEC.

Draft: feature-complete for everything verifiable on macOS; on-device bring-up
(aarch64/Sniper build, vulkan_enable2, hardware decode, real sensors, comfort)
needs the hardware. See clients/steam-frame/README.md.
…nonical

The client existed twice: this draft, and a copy outside the repo that had
picked up all the real work since. They had drifted apart in every source file.
This reconstructs the draft from the newer copy, which becomes the only one.

The reconstructed client compiles against common/protocol/include rather than a
vendored snapshot of those headers, so protocol changes reach it directly. The
vendored copy is what allowed the two to drift without anything noticing.

Brings in per-milestone work not present in the draft: foveation reconstruction
tests across every preset at real per-eye resolution, eye-gaze reporting that
drives the server's foveated-encode centre, and the eyeFov fields the server
needs to normalize gaze against the real projection instead of a default.
…ntre

The client had the whole receive half already: announce-driven foveation
parameters, per-frame centre tracking from the render pose, and the
un-warp tests. But ClientConnect went out with zero capabilities, so
the server sent a plain downscaled stream and the gaze we report drove
nothing. Advertise CLIENT_CAPABILITY_FOVEATED_ENCODING plus the new
CLIENT_CAPABILITY_FOVEATION_CENTER, which the server now requires
before it steers the centre at all.

Decode the centre through DecodeCenterShift, the shared helper the
encoder itself uses, instead of composing Dequantize + Align by hand.
…aved FEC

The client carried its own fork of the Android receiver's reassembly
state machine, third copy of the same logic, complete with the two bugs
fixed there: the almost-complete gate that deferred burst recovery to
the next frame, and parity matched by frameIndex alone across NALs that
share one. Drop the fork for the shared, tested assembler in
common/streaming; NetworkReceiver keeps sockets, NACKs, and stats.

Advertise CLIENT_CAPABILITY_FEC_INTERLEAVED and select the layout from
SERVER_FEATURE_FEC_INTERLEAVED in the announce, so the client exercises
interleaved recovery rather than only the contiguous compatibility
path. The FEC test drops its use of the removed GroupRange in favour of
GroupLayout.
… tracker

FRAME_CLIENT_SYNTHETIC_GAZE=1 substitutes a slow Lissajous sweep for
XR_EXT_eye_gaze_interaction, authored directly in head space, so the
gaze-driven foveation chain is visible end-to-end where the runtime has
no gaze extension: the sharp region glides with the sweep, pins briefly
at the horizontal extremes, and decays back to centre when the variable
is unset mid-session. Overrides real gaze when both exist so demos are
deterministic. Verifies everything except the physical eye-to-sensor
sign, which still needs eye-tracking hardware.
The centre was applied from the most recently received render pose,
but decode runs several frames deep, so the shader un-warped the
displayed frame with a centre from five or six frames ahead. Every
individual value was correct; the pairing was wrong. While the centre
moves the mismatch is centre velocity times pipeline depth, and the
periphery visibly stretches - the synthetic gaze sweep made it obvious
as continuous horizontal breathing. Exactly the per-frame identity
docs/protocol.md now spells out.

Carry the presentation timestamp through the decoder (into
av_parser_parse2 and back out of frame->pts, so parser and codec
delays cannot skew the pairing), then fetch the render pose matched to
the decoded frame via TakeRenderPoseForPresentationTimeUs and apply
pose and centre together, cached across repaints of the same texture.
The per-frame centre is only ever applied through the matched path;
the latest-received pose remains only as a timewarp fallback before
the first match. A rate-limited match hit/miss log stays in as
diagnostics - a silent fall-off in match rate is this exact bug
reappearing.

Verified over loopback with the synthetic sweep: pose match rate 100%
(360+ frames, 0 misses), the sharp region tracks the sweep across the
full frame, and peripheral cross-correlation between captures shows
only whole-image timewarp translation where it previously showed
opposite-sign stretching.
@fire

fire commented Oct 3, 2026

Copy link
Copy Markdown

Do you want me to test on the frame?

@fire

fire commented Oct 3, 2026 •

Copy link
Copy Markdown

See also my proposal to use the video codec pyrowave to drop the ffmpeg because MPL2 & LGPL3 are different licenses and also the video encode per-device licensing fee problem #45

Hardware decode (Vulkan Video / VAAPI) — replaces ~110 ms/frame software decode

pyrowave helps here.

@fire

fire commented Oct 3, 2026

Copy link
Copy Markdown

Controllers (KHR-simple + Oculus-Touch), XR_EXT_hand_tracking, XR_EXT_eye_gaze_interaction

Isn't this missing the steam frame controller profile?

@fire

fire commented Oct 3, 2026 •

Copy link
Copy Markdown

Build Linux ARM64 on the Sniper runtime; deploy via SteamOS Devkit Client

In theory you can deploy as WIN64 over proton too but that requires a Windows port of oxrsys or Android over lepton.

@demonixis
demonixis deleted the branch demonixis:develop October 3, 2026 20:13
@demonixis demonixis closed this Oct 3, 2026
@demonixis demonixis reopened this Oct 3, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants