Repository navigation
Conversation
Registers the ROCm platform attention backend "aiter_sage_attn", which runs non-causal BF16/FP16 self-attention through aiter.gfx1201_sage_attention (INT8 QK^T, FP8 PV, FP32 accumulation). The backend is explicit opt-in via attn_type and raises at construction on non-ROCm devices, on GPUs other than gfx1201, and when the installed aiter does not provide the op. Causal attention, masks, dropout and packed multi-sequence cu_seqlens are rejected. Signed-off-by: Sylvan Liu <Sylvan.Liu@amd.com>
…h aiter_sage_attn Signed-off-by: Sylvan Liu <Sylvan.Liu@amd.com>
On ROCm, SolAttnWeight calls aiter.gfx1201_sol_attention (same arguments as the NVlabs sol_attn forward: tau, diag threshold, sink tokens) instead of the CUDA sol_attn package. The ROCm path is restricted to gfx1201 and to thresh_type='diag' with the default compile mode; other GPUs report an ineligible call (strict mode raises, otherwise the existing SDPA fallback). aiter_sage_attn is accepted as dense guard backend, and a MiniMax-H3 Ref2AV config for 2 x R9700 is added. Signed-off-by: Sylvan Liu <Sylvan.Liu@amd.com>
This was referenced Oct 11, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
This PR runs Sol-Attn on AMD Radeon RDNA4 (gfx1201: Radeon AI PRO R9700 / R9600D, RX 9070 series) through aiter.
Today
sol_attnneeds the CUDA NVlabssol_attnpackage; on ROCmSolAttnWeightreports the call as ineligible andfalls back to SDPA (or raises with
strict).Changes
lightx2v/common/ops/attn/sol_attn.py, ROCm only (CUDA path unchanged):aiter.gfx1201_sol_attention, which takes the same arguments as the NVlabssol_attnforward(
tau, diag threshold, sink tokens, softmax scale);kv_splitsdoes not apply to it and is ignored;SDPA fallback);
set_configraises on ROCm forthresh_typeother thandiagandcompile_modeother thandefault;aiter_sage_attn([ROCm] Add aiter_sage_attn attention backend for gfx1201 (RDNA4) #1598) is accepted asdense_backend;configs/platforms/amd_rocm/minimax_h3_ref2av_r9700_sol.json: the [ROCm] Add aiter_sage_attn attention backend for gfx1201 (RDNA4) #1598 MiniMax-H3 config withattn_type: sol_attn(
tau1.5,diag,dense_steps6,dense_layers[0], dense backendaiter_sage_attn,strict).scripts/platforms/amd_rocm/run_minimax_h3_ref2av_r9700.sh:soloption.sol_attnis selected only byattn_type; no defaults change.Test
2 x R9700 (gfx1201, 32 GB), ROCm 7.1,
PLATFORM=amd_rocm, MiniMax-H3 ref2av TP2, inputstreet_dance(
girl.png,img_0.jpg), seed 42, 768x1344, 362 frames, DiT block offload, AdaLN cache. Only the DiT self-attentiondiffers between runs.
SolAttnWeightoutput (3D/4D, softmax scale) and the dense guard (by step, by layer) bitwise identical to directaiter.gfx1201_sol_attention/aiter.gfx1201_sage_attentioncalls; unsupportedthresh_type/compile_modeandFP32 inputs with
strictraise.ruff check/ruff format --check(v0.11.0) clean.End to end on
main(4fe984c + #1595 + #1596, 3 steps,dense_steps1,ROCM_DETERMINISTIC_FP32_BLAS=1):this PR vs
mainwithout it, with the same aiter ops installed as thesol_attn/sageattentionpackages (stockCUDA code path of
SolAttnWeight,sage_attn2dense guard):dense_backendaiter_sage_attn)main+ packagesVideo and audio bitwise identical. Per rank 98 Sol calls and 52 dense-guard calls. For comparison, dense
aiter_sage_attnonmain(#1598, other GPU pair) runs at 89.0 s/step, so a Sol step is 1.63x faster.End to end at the stock 29 steps (LightX2V 0.5.0 unmodified,
sol_attnwith thesol_attn_settingofminimax_h3_sol_block_offload.json(the settings of the config added here), the aiter ops installed as thesol_attn/sageattentionpackages, i.e. the path this PR is bitwise equal to onmain; dense guard on steps 1-6 and layer 0):gfx1201_sage_attention(code object)gfx1201_sol_attentionHIP kernelgfx1201_sol_attentioncode objectPer rank 1127 Sol calls and 323 dense-guard calls, 0 fallbacks, all outputs finite.
Over 29 steps the sampling trajectory is chaotic (reordering the FP32 sums of exact attention alone gives 14.6 dB
against the SDPA run), so pixel metrics against one reference run do not measure attention accuracy. Global statistics
of the Sol run (luma 70.0 / 56.9, sharpness 40.2, audio RMS 0.249) are close to the
dense run (69.3 / 56.9, 40.6, 0.238). Sol is block-sparse and lossy by design; its accuracy against FP32 attention
on real MiniMax-H3 tensors is reported with the aiter op.
29-step frames 0/120/241/361 (columns: BF16 SDPA, dense code object, Sol HIP kernel, Sol code object):
Before ready for review
gfx1201_sol_attentionmergedscripts/platforms/amd_rocm/run_minimax_h3_ref2av_r9700.sh solas shipped (29 steps, with the AdaLN cachestep) on the merged aiter and post the numbers here