Contributors: Hasan Unlu, Siqin Liu, Tin Nguyen, Rohit Rao, Dave Wei, Hiruna Vishwamith, Yinuo Zhao
Contact: hunlu@apexcompute.com, siqin.liu@apexcompute.com, tin.nguyen@apexcompute.com, rohit@apexcompute.com, dave.wei@apexcompute.com, hiruna@apexcompute.com, yinuo.zhao@apexcompute.com
⚙️ Hardware Architecture Update v1.41(update_87eabea5.bin)

🛒 Purchase FPGA Board with Unified Engine IP Block for $49.99
Includes ongoing hardware design updates so you always have the latest architecture.
This guide covers installation and usage of the Xilinx XDMA driver for PCIe-based FPGA communication.
- Kernel headers installed:
sudo apt install linux-headers-$(uname -r)
Clone the official Xilinx DMA driver repository:
git clone https://github.com/Xilinx/dma_ip_drivers.git
cd dma_ip_drivers/XDMA/linux-kernel/xdma
sudo make installTip: If
sudo make installfails, you may need to disable Secure Boot in your BIOS settings.
Load the XDMA driver with interrupt mode 0 (auto-detect):
sudo insmod /lib/modules/$(uname -r)/xdma/xdma.ko interrupt_mode=0Apply the following script
# 1. Remove any conflicting configs
sudo rm -f /etc/modprobe.d/blacklist-xdma.conf \
/etc/modprobe.d/xdma.conf \
/etc/modules-load.d/xdma.conf
# 2. Create systemd service
sudo tee /etc/systemd/system/xdma.service << 'EOF'
[Unit]
Description=Xilinx XDMA Driver
After=local-fs.target
[Service]
Type=oneshot
ExecStart=/bin/sh -c '/sbin/insmod /lib/modules/$(uname -r)/xdma/xdma.ko || true'
ExecStartPost=/bin/sh -c 'chmod 666 /dev/xdma*'
RemainAfterExit=yes
[Install]
WantedBy=multi-user.target
EOF
# 3. Enable and start
sudo systemctl daemon-reload
sudo systemctl enable xdma
sudo systemctl restart xdma
# 4. Verify
sudo systemctl status xdma
ls -la /dev/xdma* | head -5python3 -m venv ~/my_torch_env
source ~/my_torch_env/bin/activate
pip install -r requirements.txtpython3 user_hw_test.pyThe Gemma3 test downloads the gated google/gemma-3-1b-it model from Hugging Face. You need to:
- Create a Hugging Face account at https://huggingface.co
- Accept the Gemma license at https://huggingface.co/google/gemma-3-1b-it
- Create an access token at https://huggingface.co/settings/tokens
- Log in from the command line:
pip install huggingface-hub
huggingface-cli loginThen run:
python3 models/gemma3/gemma3_test.py --prompt "your prompt"Qwen2.5-Omni U55C: do not run this generic update procedure for the Omni runner. It validates the already-installed supported image and never reprograms the FPGA; see
models/qwen2.5_omni_7b/README.md.
The current hardware release is v1.4 (update_006e0d2f.bin). At startup the software reads the FPGA version register and checks it against the expected release hash (0x006e0d2f); on a mismatch it stops and tells you which bin to flash.
./update_fpga.sh
One command does the whole update, no PC reboot or power cycle. With no arguments it picks up the update_*.bin in the repo root, checks the running version first and exits immediately if the FPGA is already up to date; otherwise it programs and verifies the flash, warm-boots the FPGA from the new image (ICAPE2 IPROG) via update_flash.py, hot-rescans the PCIe bus (sudo ./rescan_xilinx.sh — the one step that needs root), then reads the FPGA version back and prints UPDATE SUCCESSFUL when it matches the bin.
If the image currently running predates the warm-boot block, the flash is still written but there is nothing to warm boot into: the script says so, tells you to cold reboot once, and stops before the PCIe rescan (exit code 3). After that one cold boot every update is reboot-free.
Options:
./update_fpga.sh --bin update_006e0d2f.bin # explicit image (or pass it positionally)
./update_fpga.sh --check # device ID + running FPGA hash vs the repo bin
./update_fpga.sh --boot # no reflash: warm boot from flash, rescan
./update_fpga.sh --force # reflash even if already up to date
provision_kintex7.py handles the unified-engine
XC7K480T + MT28GU512 BPI-x16, 64 MiB board. Source Vivado's
settings64.sh first; Python 3.8+ and Vivado Hardware Manager are required.
It works independently of the PCIe/XDMA device numbering.
# Inventory all cables and devices; no programming.
python3 provision_kintex7.py --list
# Read-only preflight (also the default without --check).
python3 provision_kintex7.py encrypted.bin --key /secure/andromeda_wrapper.nky --check
# Permanently burn the key, protect it, then erase/program/verify BPI flash.
python3 provision_kintex7.py encrypted.bin --key /secure/andromeda_wrapper.nky --program
# Optional: pin the cable/device/DNA printed by --list or --check, and boot.
python3 provision_kintex7.py encrypted.bin --key /secure/andromeda_wrapper.nky \
--target CABLE_SERIAL --device xc7k480t_0 --dna DEVICE_DNA --program --boot
# Subsequent updates: use an image encrypted for the already-fused key.
python3 provision_kintex7.py encrypted.bin --flash-only --boot
# Make equivalent: omit PROGRAM=1 for the read-only check.
make kintex7_efuse_flash BINFILE=encrypted.bin NKY_FILE=/secure/andromeda_wrapper.nky PROGRAM=1The matching .nky is required for provisioning: the AES key cannot be
extracted from an encrypted .bin. If --key is omitted, encrypted.bin
uses encrypted.nky beside it. Keep both files from the same build. The
input must be the encrypted flash image generated with
write_cfgmem -format bin -size 64 -interface BPIx16 -loadbit "up 0x0 design.bit";
the design must use BITSTREAM.ENCRYPTION.ENCRYPTKEYSELECT EFUSE.
The script checks the BIN's clear encryption header and payload length and
the NKY's device/key format. These checks do not prove the key matches the
ciphertext, authenticate the image, or identify the part inside its encrypted
payload. Use the XC7K480T build's matching image/key pair.
Every cable is scanned unless --target specifies an exact target path or
unique serial. Selection must produce exactly one XC7K480T. Unreachable
cables, failed device reads and ambiguous matches stop the operation.
The device is identified by cable, device name, part and DNA, then reacquired
and checked after reopening its cable. No fallback selects the first FPGA.
--server HOST:3121 selects a remote hardware server.
--program is irreversible authorization. It refuses an already-fused
AES key or locked key/control registers. It uses the existing Andromeda policy:
FUSE_USER=0 and FUSE_CNTL=0x0c (key write/read protection). Programming
the AES key consumes the opportunity to provision FUSE_USER[7:0]; this
flow fixes those bits at zero. It leaves CFG_AES_Only unset because setting
that bit prevents Vivado indirect BPI flash programming. This flow does not
enforce encrypted-only configuration. See AMD's
7-series encryption application note, XAPP1239.
Use an eFUSE-capable JTAG cable
and a powered board with stable supply rails.
Programming replaces the running FPGA configuration with Vivado's flash helper,
so stop workloads first. The script checks AES-programmed status and protection
bits after burning, then programs flash with verification enabled. --boot
boots from flash and checks DONE; without it, power-cycle the board to load
the image. PCIe rescanning is separate. --flash-only never burns fuses and
requires an already-programmed AES key; it cannot compare a read-protected key
with the input image. A verified flash write alone does not prove successful
decryption or application operation.
Each run keeps a device record, result JSON and any generated secret NKZ
export in a private directory under ~/.local/state/unified-engine/efuse/
(override with --output-dir). Input snapshots are deleted on exit; Vivado
logs/journals are disabled and long key values are redacted from its console
output. Protect NKZ exports like the original key. If flash fails after the
fuse burn succeeds, preserve the export and retry with --flash-only using
the same encrypted image. Never attempt to burn a replacement key.
Hardware-independent regression checks:
python3 -m unittest discover -s tests -p 'test_provision_kintex7.py'Gemma3 above is just the quick-start example. Every model below runs on the
engine today; each folder has its own README/config, and most LLMs ship a
*_run_from_bin.py for execute-only deploys from precompiled bins.
| Model | Folder | Type |
|---|---|---|
| Gemma 3 1B | models/gemma3 |
Text LM |
| Gemma 4 E2B | models/gemma4_e2b |
Multimodal LM (text, vision, audio) |
| Gemma 4 E4B | models/gemma4_e4b |
Multimodal LM (text, vision, audio) |
| Llama 3.2 1B | models/llama3.2_1b |
Text LM |
| Llama 3.2 3B | models/llama3.2_3b |
Text LM |
| Qwen3 0.6B | models/qwen3_0.6b |
Text LM |
| Qwen3 1.7B | models/qwen3_1.7b |
Text LM |
| Qwen3 4B | models/qwen3_4b |
Text LM |
| Qwen3.5 2B | models/qwen3.5_2b |
Text LM |
| Qwen2.5-VL 3B | models/qwen2.5_vl_3b |
Vision-language |
| Qwen2.5-Omni 7B | models/qwen2.5_omni_7b |
Multimodal Thinker (text, image, audio -> text) |
| SmolVLM2 | models/smolvlm2 |
Vision-language |
| GPT-2 | models/gpt2 |
Text LM |
| LocateAnything 3B | models/locateanything_3b |
Open-vocabulary localization |
| MobileNetV2 (224 + SSD-FPNLite 640) | models/mobilenetv2 |
Classification / detection |
| Parakeet | models/parakeet |
Speech recognition (incl. streaming) |
| MobileSAM | models/mobilesam |
Segmentation |
| Swin | models/swin |
Image classification |
Qwen2.5-Omni-7B currently accelerates the Thinker path: text, image, and audio inputs produce text. It targets the 8 GiB Alveo U55 configuration and runs on engines 0-7; HW_INFO must report 8 GiB and at least eight available engines.
# Text
python models/qwen2.5_omni_7b/qwen2.5_omni_7b_test.py --multi-core 8 \
--prompt "If x + 3 = 5, what is x?"
# Image (bare --image uses test_samples/yosemite.jpg)
python models/qwen2.5_omni_7b/qwen2.5_omni_7b_test.py --multi-core 8 --image
# Audio
python models/qwen2.5_omni_7b/qwen2.5_omni_7b_test.py --multi-core 8 \
--audio test_samples/apex.wav --prompt "Transcribe the speech exactly."Run the whole suite (or a subset) with the automated tester:
make model_test run_from_bin # skip the pre-clean, reuse existing compiled bins
make model_test gemma4_e2b # one model
make model_test_help # all modesNotes on the two modes:
- Without the
run_from_binword,model_testrunsmake cleanfirst, which deletes cached model bins and rebuilds everything from the HF models (slow; needs the HF models available). run_from_binskips the pre-clean so models with a*_run_from_bin.pyruntime (the LLM/VLM rows above) reuse their bins. gemma3, gpt2 and the vision/speech models have no runtime-only entry yet and still run through their*_test.pyscripts, which need their model assets on disk. On a deploy host that only has pregenerated bins, run the runtime-only subset by name, e.g.make model_test gemma4_e2b llama3.2_1b qwen3_4b run_from_bin.
All benchmarks were collected on RTL running on a Kintex UltraScale+ FPGA in real time.
📄 Download Benchmark Datasheet (PDF)
| Specification | Value |
|---|---|
| Engine frequency | 333 MHz |
| Theoretical peak (BF16) | 42 GFLOPS/s |
| Memory interface | DDR4 @ 1333 MHz, 32-bit |
| AXI Master Data Width | 256 bits |
| On-chip SRAM | 1.05 MB |
| Total power | 4.5 W |
| BF16 MatMul | 40.17 GFLOPS/s (95.6% utilization) |
| BF16 MatMul + Bias + Activation | 40.03 GFLOPS/s (95.3% utilization) |
| BF16 Softmax MatMul | 37.76 GFLOPS/s (89.9% utilization) |
| Memory-Efficient Attention | ~90% utilization |
| Quantized MatMul (BF16 × INT4/FP4) | 40.03 GFLOPS/s (95.3% utilization) |
| Quantized MatVec (Streaming matrix, decoding mode friendly) (BF16 × INT4/FP4) | 31.33 GFLOPS/s (74.6% utilization) |
| RMSNorm | 4.81 GFLOPS/s |
| LayerNorm | 5.90 GFLOPS/s |
| Quantize (BF16 → INT4/FP4) | 5.72 GFLOPS/s |
| Dequantize (INT4/FP4 → BF16) | 3.31 GFLOPS/s |
| Hardware trace buffer | 8,192 timestamps |
| Multi-engine tensor parallelism | Supported with Synchronization Flag instructions |
| Parameter | Value |
|---|---|
| Memory interface | DDR4 at 1333 MHz, 32-bit data path |
| Engine frequency | 333 MHz |
| Memory interface clock | Synchronized 1:1 with engine clock |
| Data width | 256 bits |
| Total power consumption | 4.5 W |
| Total on-chip SRAM | 1.05 MB |
Total floating-point operations per second from the engine at 333 MHz is approximately 42 GFLOPS/s.
| Name | CLB LUTs | CLB Registers | Block RAM Tile | URAM | DSPs |
|---|---|---|---|---|---|
| unified_engine_top | 78,348 | 50,045 | 16 | 30 | 197 |
| Operation | FLOPS |
|---|---|
| FMA (Fused Multiply-Add) | 2 |
| Addition / Multiplication | 1 |
| Exponent | 1 |
| Division | 1 |
Engine speed: 333 MHz; theoretical peak: 42 GFLOPS/s. Metrics based on M=1024, K=1024, N=1024. O denotes the output tensor. All matrix-matrix operations we are reaching up to 95% FLOPS utilizations.
| Op | Operands | FLOPS | Cycles (latency) | Achieved GFLOPS/s |
|---|---|---|---|---|
| A Bᵀ | A[M,K], B[N,K] → O[M,N] | 2MKN | 17,820,455 (53.3 ms) | 40.17 |
| A Bᵀ + C | A[M,K], B[N,K], C[M,N] → O[M,N] | 2MKN + MN | 17,858,564 (53.5 ms) | 40.10 |
| GELU(A Bᵀ) | A[M,K], B[N,K] → O[M,N] | 2MKN + 4MN | 17,923,045 (53.7 ms) | 40.02 |
| GELU(A Bᵀ + C) | A[M,K], B[N,K], C[M,N] → O[M,N] | 2MKN + MN + 4MN | 17,927,850 (53.7 ms) | 40.03 |
| SiLU(A Bᵀ) | A[M,K], B[N,K] → O[M,N] | 2MKN + 4MN | 17,921,594 (53.7 ms) | 40.02 |
| SiLU(A Bᵀ + C) | A[M,K], B[N,K], C[M,N] → O[M,N] | 2MKN + MN + 4MN | 17,926,623 (53.7 ms) | 40.03 |
| softmax(A Bᵀ) | A[M,K], B[N,K] → O[M,N] | 2MKN + 5MN | 19,004,997 (57.01 ms) | 37.76 |
| softmax(A Bᵀ + C) | A[M,K], B[N,K], C[M,N] → O[M,N] | 2MKN + MN + 5MN | 19,051,310 (57.15 ms) | 37.68 |
| Aᵀ | A[M,N] → O[N,M] | 0 | 1,648,647 (4.9 ms) | N/A |
| A · scalar | A[M,N] → O[M,N] | MN | 180,500 (541 µs) | 1.94 |
| A + scalar | A[M,N] → O[M,N] | MN | 181,005 (543 µs) | 1.93 |
| A · B | A[M,N], B[M,N] → O[M,N] | MN | 263,580 (790 µs) | 1.33 |
| A + B | A[M,N], B[M,N] → O[M,N] | MN | 263,871 (791 µs) | 1.33 |
| RMSNorm(A) · γ | A[M,N], γ[N] → O[M,N] | 4MN | 290,945 (872 µs) | 4.81 |
| LayerNorm(A) · γ + β | A[M,N], γ[N], β[N] → O[M,N] | 7MN | 414,679 (1.24 ms) | 5.90 |
The following kernel computes the attention block for given query/key/value tensors and an optional mask or bias. It reaches almost 90% utilization of theoretical FLOPS.
memory_efficient_attention(q, k, v, mask_or_bias)
Equivalent PyTorch reference:
def memory_efficient_attention(q, k, v, attn_bias=None):
scale = 1.0 / math.sqrt(head_dim)
attn_weights = (q @ k.T) * scale
if attn_bias is not None:
attn_weights = attn_weights + attn_bias
scores = torch.softmax(attn_weights, dim=-1)
return scores @ v

Flash attention benchmark — bias off

Flash attention benchmark — bias on
Engine speed: 333 MHz; theoretical peak: 42 GFLOPS/s. In quantized mode, achieved FLOPS are the same for any M. In contrast, for tiled matrix-matrix multiplication, smaller M reduces FLOPS utilization. fp4 refers to nvfp4 (Nvidia fp4).
Metrics based on M=1024, K=1024, N=1024.
| Op | Precision | Operands | FLOPS | Cycles (latency) | Achieved GFLOPS/s |
|---|---|---|---|---|---|
| A Bᵀ | A(bf16) B(int4/fp4) O(bf16) | A[M,K], B[N,K] → O[M,N] | 2MKN | 22,849,177 (68.5 ms) | 31.33 |
| A Bᵀ + C | A(bf16) B(int4/fp4) C(bf16) O(bf16) | A[M,K], B[N,K], C[M,N] → O[M,N] | 2MKN + MN | 23,073,635 (69.2 ms) | 31.04 |
| GELU(A Bᵀ) | A(bf16) B(int4/fp4) O(bf16) | A[M,K], B[N,K] → O[M,N] | 2MKN + 4MN | 22,850,336 (68.5 ms) | 31.39 |
| GELU(A Bᵀ + C) | A(bf16) B(int4/fp4) C(bf16) O(bf16) | A[M,K], B[N,K], C[M,N] → O[M,N] | 2MKN + MN + 4MN | 23,100,231 (69.3 ms) | 31.06 |
| SiLU(A Bᵀ) | A(bf16) B(int4/fp4) O(bf16) | A[M,K], B[N,K] → O[M,N] | 2MKN + 4MN | 22,850,243 (68.5 ms) | 31.39 |
| SiLU(A Bᵀ + C) | A(bf16) B(int4/fp4) C(bf16) O(bf16) | A[M,K], B[N,K], C[M,N] → O[M,N] | 2MKN + MN + 4MN | 23,104,094 (69.3 ms) | 31.06 |
| Op | Precision | Operands | FLOPS | Cycles (latency) | Achieved GFLOPS/s |
|---|---|---|---|---|---|
| Quantize(A) | A(bf16) O(int4/fp4) | A[N] → O[N] | 2N | 15,266 (45.8 µs) | 5.72 |
| Dequantize(A) | A(int4/fp4) O(bf16) | A[N] → O[N] | N | 13,193 (39.5 µs) | 3.31 |
The engine includes a hardware trace buffer capable of recording 8,192 timestamps, allowing cycle-accurate profiling of kernel execution. This is useful for experimenting with tensor parallelism across multiple engines.
The example below demonstrates splitting a 256×2048 @ 2048×1024 matrix multiplication across two engines:
- Engine 0: 192×2048 @ 2048×1024 (larger partition)
- Engine 1: 64×2048 @ 2048×1024 (smaller partition)
Because the two partitions have unequal workloads, the smaller partition finishes before the larger one. A hardware synchronization flag is used to hold the faster engine until both are complete before proceeding to the next stage. The trace visualization below shows this synchronization in action — the idle gap on Engine 1 is where it waits for Engine 0 to finish.

Trace buffer visualization — 256×2048 @ 2048×1024 split across two engines with hardware synchronization
