Skip to content

About

No description, website, or topics provided.

Resources

Stars

96 stars

Watchers

0 watching

Forks

Latest commit

 

History

886 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Apex Compute

Contributors: Hasan Unlu, Siqin Liu, Tin Nguyen, Rohit Rao, Dave Wei, Hiruna Vishwamith, Yinuo Zhao

Contact: hunlu@apexcompute.com, siqin.liu@apexcompute.com, tin.nguyen@apexcompute.com, rohit@apexcompute.com, dave.wei@apexcompute.com, hiruna@apexcompute.com, yinuo.zhao@apexcompute.com

⚙️ Hardware Architecture Update v1.41(update_87eabea5.bin)

FPGA Board
🛒 Purchase FPGA Board with Unified Engine IP Block for $49.99
Includes ongoing hardware design updates so you always have the latest architecture.

Discord

XDMA Driver Setup and Usage Guide

This guide covers installation and usage of the Xilinx XDMA driver for PCIe-based FPGA communication.

Prerequisites

  • Kernel headers installed: sudo apt install linux-headers-$(uname -r)

Installation

1. Install XDMA Driver from Xilinx Repository

Clone the official Xilinx DMA driver repository:

git clone https://github.com/Xilinx/dma_ip_drivers.git
cd dma_ip_drivers/XDMA/linux-kernel/xdma
sudo make install

Tip: If sudo make install fails, you may need to disable Secure Boot in your BIOS settings.

2. Load the Driver

Load the XDMA driver with interrupt mode 0 (auto-detect):

sudo insmod /lib/modules/$(uname -r)/xdma/xdma.ko interrupt_mode=0

3. Load the Driver Every Boot Automatically (Recommended)

Apply the following script

# 1. Remove any conflicting configs
sudo rm -f /etc/modprobe.d/blacklist-xdma.conf \
           /etc/modprobe.d/xdma.conf \
           /etc/modules-load.d/xdma.conf

# 2. Create systemd service
sudo tee /etc/systemd/system/xdma.service << 'EOF'
[Unit]
Description=Xilinx XDMA Driver
After=local-fs.target

[Service]
Type=oneshot
ExecStart=/bin/sh -c '/sbin/insmod /lib/modules/$(uname -r)/xdma/xdma.ko || true'
ExecStartPost=/bin/sh -c 'chmod 666 /dev/xdma*'
RemainAfterExit=yes

[Install]
WantedBy=multi-user.target
EOF

# 3. Enable and start
sudo systemctl daemon-reload
sudo systemctl enable xdma
sudo systemctl restart xdma

# 4. Verify
sudo systemctl status xdma
ls -la /dev/xdma* | head -5

4. Set Up Python Environment

python3 -m venv ~/my_torch_env
source ~/my_torch_env/bin/activate
pip install -r requirements.txt

5. Run Hardware Tests

python3 user_hw_test.py

6. Run Gemma3 Inference (requires Hugging Face)

The Gemma3 test downloads the gated google/gemma-3-1b-it model from Hugging Face. You need to:

  1. Create a Hugging Face account at https://huggingface.co
  2. Accept the Gemma license at https://huggingface.co/google/gemma-3-1b-it
  3. Create an access token at https://huggingface.co/settings/tokens
  4. Log in from the command line:
pip install huggingface-hub
huggingface-cli login

Then run:

python3 models/gemma3/gemma3_test.py --prompt "your prompt"

7. Updating HW bin file

Qwen2.5-Omni U55C: do not run this generic update procedure for the Omni runner. It validates the already-installed supported image and never reprograms the FPGA; see models/qwen2.5_omni_7b/README.md.

The current hardware release is v1.4 (update_006e0d2f.bin). At startup the software reads the FPGA version register and checks it against the expected release hash (0x006e0d2f); on a mismatch it stops and tells you which bin to flash.

./update_fpga.sh

One command does the whole update, no PC reboot or power cycle. With no arguments it picks up the update_*.bin in the repo root, checks the running version first and exits immediately if the FPGA is already up to date; otherwise it programs and verifies the flash, warm-boots the FPGA from the new image (ICAPE2 IPROG) via update_flash.py, hot-rescans the PCIe bus (sudo ./rescan_xilinx.sh — the one step that needs root), then reads the FPGA version back and prints UPDATE SUCCESSFUL when it matches the bin.

If the image currently running predates the warm-boot block, the flash is still written but there is nothing to warm boot into: the script says so, tells you to cold reboot once, and stops before the PCIe rescan (exit code 3). After that one cold boot every update is reboot-free.

Options:

./update_fpga.sh --bin update_006e0d2f.bin   # explicit image (or pass it positionally)
./update_fpga.sh --check                     # device ID + running FPGA hash vs the repo bin
./update_fpga.sh --boot                      # no reflash: warm boot from flash, rescan
./update_fpga.sh --force                     # reflash even if already up to date

Kintex-7 encrypted BIN + eFUSE provisioning over JTAG

provision_kintex7.py handles the unified-engine XC7K480T + MT28GU512 BPI-x16, 64 MiB board. Source Vivado's settings64.sh first; Python 3.8+ and Vivado Hardware Manager are required. It works independently of the PCIe/XDMA device numbering.

# Inventory all cables and devices; no programming.
python3 provision_kintex7.py --list

# Read-only preflight (also the default without --check).
python3 provision_kintex7.py encrypted.bin --key /secure/andromeda_wrapper.nky --check

# Permanently burn the key, protect it, then erase/program/verify BPI flash.
python3 provision_kintex7.py encrypted.bin --key /secure/andromeda_wrapper.nky --program

# Optional: pin the cable/device/DNA printed by --list or --check, and boot.
python3 provision_kintex7.py encrypted.bin --key /secure/andromeda_wrapper.nky \
  --target CABLE_SERIAL --device xc7k480t_0 --dna DEVICE_DNA --program --boot

# Subsequent updates: use an image encrypted for the already-fused key.
python3 provision_kintex7.py encrypted.bin --flash-only --boot

# Make equivalent: omit PROGRAM=1 for the read-only check.
make kintex7_efuse_flash BINFILE=encrypted.bin NKY_FILE=/secure/andromeda_wrapper.nky PROGRAM=1

The matching .nky is required for provisioning: the AES key cannot be extracted from an encrypted .bin. If --key is omitted, encrypted.bin uses encrypted.nky beside it. Keep both files from the same build. The input must be the encrypted flash image generated with write_cfgmem -format bin -size 64 -interface BPIx16 -loadbit "up 0x0 design.bit"; the design must use BITSTREAM.ENCRYPTION.ENCRYPTKEYSELECT EFUSE. The script checks the BIN's clear encryption header and payload length and the NKY's device/key format. These checks do not prove the key matches the ciphertext, authenticate the image, or identify the part inside its encrypted payload. Use the XC7K480T build's matching image/key pair.

Every cable is scanned unless --target specifies an exact target path or unique serial. Selection must produce exactly one XC7K480T. Unreachable cables, failed device reads and ambiguous matches stop the operation. The device is identified by cable, device name, part and DNA, then reacquired and checked after reopening its cable. No fallback selects the first FPGA. --server HOST:3121 selects a remote hardware server.

--program is irreversible authorization. It refuses an already-fused AES key or locked key/control registers. It uses the existing Andromeda policy: FUSE_USER=0 and FUSE_CNTL=0x0c (key write/read protection). Programming the AES key consumes the opportunity to provision FUSE_USER[7:0]; this flow fixes those bits at zero. It leaves CFG_AES_Only unset because setting that bit prevents Vivado indirect BPI flash programming. This flow does not enforce encrypted-only configuration. See AMD's 7-series encryption application note, XAPP1239. Use an eFUSE-capable JTAG cable and a powered board with stable supply rails.

Programming replaces the running FPGA configuration with Vivado's flash helper, so stop workloads first. The script checks AES-programmed status and protection bits after burning, then programs flash with verification enabled. --boot boots from flash and checks DONE; without it, power-cycle the board to load the image. PCIe rescanning is separate. --flash-only never burns fuses and requires an already-programmed AES key; it cannot compare a read-protected key with the input image. A verified flash write alone does not prove successful decryption or application operation.

Each run keeps a device record, result JSON and any generated secret NKZ export in a private directory under ~/.local/state/unified-engine/efuse/ (override with --output-dir). Input snapshots are deleted on exit; Vivado logs/journals are disabled and long key values are redacted from its console output. Protect NKZ exports like the original key. If flash fails after the fuse burn succeeds, preserve the export and retry with --flash-only using the same encrypted image. Never attempt to burn a replacement key.

Hardware-independent regression checks:

python3 -m unittest discover -s tests -p 'test_provision_kintex7.py'

Supported Models

Gemma3 above is just the quick-start example. Every model below runs on the engine today; each folder has its own README/config, and most LLMs ship a *_run_from_bin.py for execute-only deploys from precompiled bins.

Model Folder Type
Gemma 3 1B models/gemma3 Text LM
Gemma 4 E2B models/gemma4_e2b Multimodal LM (text, vision, audio)
Gemma 4 E4B models/gemma4_e4b Multimodal LM (text, vision, audio)
Llama 3.2 1B models/llama3.2_1b Text LM
Llama 3.2 3B models/llama3.2_3b Text LM
Qwen3 0.6B models/qwen3_0.6b Text LM
Qwen3 1.7B models/qwen3_1.7b Text LM
Qwen3 4B models/qwen3_4b Text LM
Qwen3.5 2B models/qwen3.5_2b Text LM
Qwen2.5-VL 3B models/qwen2.5_vl_3b Vision-language
Qwen2.5-Omni 7B models/qwen2.5_omni_7b Multimodal Thinker (text, image, audio -> text)
SmolVLM2 models/smolvlm2 Vision-language
GPT-2 models/gpt2 Text LM
LocateAnything 3B models/locateanything_3b Open-vocabulary localization
MobileNetV2 (224 + SSD-FPNLite 640) models/mobilenetv2 Classification / detection
Parakeet models/parakeet Speech recognition (incl. streaming)
MobileSAM models/mobilesam Segmentation
Swin models/swin Image classification

Qwen2.5-Omni-7B currently accelerates the Thinker path: text, image, and audio inputs produce text. It targets the 8 GiB Alveo U55 configuration and runs on engines 0-7; HW_INFO must report 8 GiB and at least eight available engines.

# Text
python models/qwen2.5_omni_7b/qwen2.5_omni_7b_test.py --multi-core 8 \
  --prompt "If x + 3 = 5, what is x?"

# Image (bare --image uses test_samples/yosemite.jpg)
python models/qwen2.5_omni_7b/qwen2.5_omni_7b_test.py --multi-core 8 --image

# Audio
python models/qwen2.5_omni_7b/qwen2.5_omni_7b_test.py --multi-core 8 \
  --audio test_samples/apex.wav --prompt "Transcribe the speech exactly."

Run the whole suite (or a subset) with the automated tester:

make model_test run_from_bin   # skip the pre-clean, reuse existing compiled bins
make model_test gemma4_e2b     # one model
make model_test_help           # all modes

Notes on the two modes:

  • Without the run_from_bin word, model_test runs make clean first, which deletes cached model bins and rebuilds everything from the HF models (slow; needs the HF models available).
  • run_from_bin skips the pre-clean so models with a *_run_from_bin.py runtime (the LLM/VLM rows above) reuse their bins. gemma3, gpt2 and the vision/speech models have no runtime-only entry yet and still run through their *_test.py scripts, which need their model assets on disk. On a deploy host that only has pregenerated bins, run the runtime-only subset by name, e.g. make model_test gemma4_e2b llama3.2_1b qwen3_4b run_from_bin.

Apex Compute Unified Engine v1.1 — Benchmark Results

All benchmarks were collected on RTL running on a Kintex UltraScale+ FPGA in real time.

Benchmark Datasheet

📄 Download Benchmark Datasheet (PDF)

Specification Value
Engine frequency 333 MHz
Theoretical peak (BF16) 42 GFLOPS/s
Memory interface DDR4 @ 1333 MHz, 32-bit
AXI Master Data Width 256 bits
On-chip SRAM 1.05 MB
Total power 4.5 W
BF16 MatMul 40.17 GFLOPS/s (95.6% utilization)
BF16 MatMul + Bias + Activation 40.03 GFLOPS/s (95.3% utilization)
BF16 Softmax MatMul 37.76 GFLOPS/s (89.9% utilization)
Memory-Efficient Attention ~90% utilization
Quantized MatMul (BF16 × INT4/FP4) 40.03 GFLOPS/s (95.3% utilization)
Quantized MatVec (Streaming matrix, decoding mode friendly) (BF16 × INT4/FP4) 31.33 GFLOPS/s (74.6% utilization)
RMSNorm 4.81 GFLOPS/s
LayerNorm 5.90 GFLOPS/s
Quantize (BF16 → INT4/FP4) 5.72 GFLOPS/s
Dequantize (INT4/FP4 → BF16) 3.31 GFLOPS/s
Hardware trace buffer 8,192 timestamps
Multi-engine tensor parallelism Supported with Synchronization Flag instructions

FPGA Presilicon Prototype Setup

System Parameters

Parameter Value
Memory interface DDR4 at 1333 MHz, 32-bit data path
Engine frequency 333 MHz
Memory interface clock Synchronized 1:1 with engine clock
Data width 256 bits
Total power consumption 4.5 W
Total on-chip SRAM 1.05 MB

Peak Operation Rate

Total floating-point operations per second from the engine at 333 MHz is approximately 42 GFLOPS/s.

FPGA Resource Utilization

Name CLB LUTs CLB Registers Block RAM Tile URAM DSPs
unified_engine_top 78,348 50,045 16 30 197

FLOPS Definitions

Operation FLOPS
FMA (Fused Multiply-Add) 2
Addition / Multiplication 1
Exponent 1
Division 1

BF16 Operation Benchmarks

Engine speed: 333 MHz; theoretical peak: 42 GFLOPS/s. Metrics based on M=1024, K=1024, N=1024. O denotes the output tensor. All matrix-matrix operations we are reaching up to 95% FLOPS utilizations.

OpOperandsFLOPSCycles (latency)Achieved GFLOPS/s
A BᵀA[M,K], B[N,K] → O[M,N]2MKN17,820,455 (53.3 ms)40.17
A Bᵀ + CA[M,K], B[N,K], C[M,N] → O[M,N]2MKN + MN17,858,564 (53.5 ms)40.10
GELU(A Bᵀ)A[M,K], B[N,K] → O[M,N]2MKN + 4MN17,923,045 (53.7 ms)40.02
GELU(A Bᵀ + C)A[M,K], B[N,K], C[M,N] → O[M,N]2MKN + MN + 4MN17,927,850 (53.7 ms)40.03
SiLU(A Bᵀ)A[M,K], B[N,K] → O[M,N]2MKN + 4MN17,921,594 (53.7 ms)40.02
SiLU(A Bᵀ + C)A[M,K], B[N,K], C[M,N] → O[M,N]2MKN + MN + 4MN17,926,623 (53.7 ms)40.03
softmax(A Bᵀ)A[M,K], B[N,K] → O[M,N]2MKN + 5MN19,004,997 (57.01 ms)37.76
softmax(A Bᵀ + C)A[M,K], B[N,K], C[M,N] → O[M,N]2MKN + MN + 5MN19,051,310 (57.15 ms)37.68
AᵀA[M,N] → O[N,M]01,648,647 (4.9 ms)N/A
A · scalarA[M,N] → O[M,N]MN180,500 (541 µs)1.94
A + scalarA[M,N] → O[M,N]MN181,005 (543 µs)1.93
A · BA[M,N], B[M,N] → O[M,N]MN263,580 (790 µs)1.33
A + BA[M,N], B[M,N] → O[M,N]MN263,871 (791 µs)1.33
RMSNorm(A) · γA[M,N], γ[N] → O[M,N]4MN290,945 (872 µs)4.81
LayerNorm(A) · γ + βA[M,N], γ[N], β[N] → O[M,N]7MN414,679 (1.24 ms)5.90

Memory-Efficient Attention

The following kernel computes the attention block for given query/key/value tensors and an optional mask or bias. It reaches almost 90% utilization of theoretical FLOPS.

memory_efficient_attention(q, k, v, mask_or_bias)

Equivalent PyTorch reference:

def memory_efficient_attention(q, k, v, attn_bias=None):
    scale = 1.0 / math.sqrt(head_dim)
    attn_weights = (q @ k.T) * scale
    if attn_bias is not None:
        attn_weights = attn_weights + attn_bias
    scores = torch.softmax(attn_weights, dim=-1)
    return scores @ v

Flash attention benchmark (bias off)
Flash attention benchmark — bias off

Flash attention benchmark (bias on)
Flash attention benchmark — bias on

Quantized Operation Benchmarks

Engine speed: 333 MHz; theoretical peak: 42 GFLOPS/s. In quantized mode, achieved FLOPS are the same for any M. In contrast, for tiled matrix-matrix multiplication, smaller M reduces FLOPS utilization. fp4 refers to nvfp4 (Nvidia fp4).

Metrics based on M=1024, K=1024, N=1024.

OpPrecisionOperandsFLOPSCycles (latency)Achieved GFLOPS/s
A BᵀA(bf16) B(int4/fp4) O(bf16)A[M,K], B[N,K] → O[M,N]2MKN22,849,177 (68.5 ms)31.33
A Bᵀ + CA(bf16) B(int4/fp4) C(bf16) O(bf16)A[M,K], B[N,K], C[M,N] → O[M,N]2MKN + MN23,073,635 (69.2 ms)31.04
GELU(A Bᵀ)A(bf16) B(int4/fp4) O(bf16)A[M,K], B[N,K] → O[M,N]2MKN + 4MN22,850,336 (68.5 ms)31.39
GELU(A Bᵀ + C)A(bf16) B(int4/fp4) C(bf16) O(bf16)A[M,K], B[N,K], C[M,N] → O[M,N]2MKN + MN + 4MN23,100,231 (69.3 ms)31.06
SiLU(A Bᵀ)A(bf16) B(int4/fp4) O(bf16)A[M,K], B[N,K] → O[M,N]2MKN + 4MN22,850,243 (68.5 ms)31.39
SiLU(A Bᵀ + C)A(bf16) B(int4/fp4) C(bf16) O(bf16)A[M,K], B[N,K], C[M,N] → O[M,N]2MKN + MN + 4MN23,104,094 (69.3 ms)31.06

Quantization / Dequantization (N=131,072)

OpPrecisionOperandsFLOPSCycles (latency)Achieved GFLOPS/s
Quantize(A)A(bf16) O(int4/fp4)A[N] → O[N]2N15,266 (45.8 µs)5.72
Dequantize(A)A(int4/fp4) O(bf16)A[N] → O[N]N13,193 (39.5 µs)3.31

Trace Buffer and Tensor Parallelism

The engine includes a hardware trace buffer capable of recording 8,192 timestamps, allowing cycle-accurate profiling of kernel execution. This is useful for experimenting with tensor parallelism across multiple engines.

The example below demonstrates splitting a 256×2048 @ 2048×1024 matrix multiplication across two engines:

  • Engine 0: 192×2048 @ 2048×1024 (larger partition)
  • Engine 1: 64×2048 @ 2048×1024 (smaller partition)

Because the two partitions have unequal workloads, the smaller partition finishes before the larger one. A hardware synchronization flag is used to hold the faster engine until both are complete before proceeding to the next stage. The trace visualization below shows this synchronization in action — the idle gap on Engine 1 is where it waits for Engine 0 to finish.

Trace buffer visualization of tensor-parallel matrix multiplication with synchronization
Trace buffer visualization — 256×2048 @ 2048×1024 split across two engines with hardware synchronization

About

No description, website, or topics provided.

Resources

Stars

96 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages