Skip to content

CI run: Highway CPU kernels on Linux arm64 (not for merge) - #2

Draft
rcfa wants to merge 38 commits into
ci-base-t167p2from
t167p2-arm64
Draft

rcfa wants to merge 38 commits into
ci-base-t167p2from
t167p2-arm64

Conversation

@rcfa

@rcfa rcfa commented Oct 1, 2026

Copy link
Copy Markdown
Owner

Runs vmlx-swift's CI in this fork for plan 2 of T-167 phase 2: Highway on Linux arm64, a Linux leg without Highway, and an advisory ThreadSanitizer job. Not for merge; the upstream PR follows osaurus-ai/mlx#18.

🤖 Generated with Claude Code

rcfa added 30 commits September 27, 2026 17:15
scripts/check-macos-identity.py rebuilds two commits' trees from git and
compares the sources SwiftPM compiles for Cmlx on macOS, each preprocessed
under the manifest's macOS settings, and the files the Xcode project
compiles from its mlx and mlx-c folders. vmlx's CMake build is outside it.
Move Source/Cmlx/mlx to the osaurus-ai/mlx pull request that adopts
ml-explore/mlx#3019's Highway kernels, and compile them on Linux x86-64:
MLX_USE_HIGHWAY_KERNELS, the Highway runtime through wrappers in
highway-runtime/, and the submodule's headers. macOS excludes every new
source, in SwiftPM and in the Xcode project, so both compile what they
did; vmlx's CMake build compiles mlx's precision.cpp, the int8 switch,
everywhere, macOS included. Linux arm64 compiles the new sources as empty
files, all but the int8 switch, until the kernels are tested there.
VMLX_HWY_ALL_TARGETS=1 adds HWY_COMPILE_ALL_ATTAINABLE for the kernel
tests. vmlx's CMake build takes Highway from the submodule.
Swift reads and sets them through it: the Highway targets, the int8
quantized-matmul switch, the thread pool and the per-family dispatch
counters. Builds without Highway kernels answer as such a build does: no
targets, int8 off, one thread, no counts.
The CPU backend's kernel tests run on Linux under one lock, since the Highway
target mask, the int8 switch and the dispatch counters are process-wide.
scripts/run-cpu-kernel-tests.sh builds the all-targets test build and runs
them twice, with MLX_CPU_THREADS=1 and with the default pool. The targets a
host must execute come from Highway's own detection of its CPU, intersected
with the targets Highway attains there, never from what the build contains.
The runner's failure excerpt printed swift-testing's summary lines but not
the values it prints under each recorded issue, so a failure showed that an
expectation failed, not with what values. Print each recorded issue with
its values, and any stall line.
ULPBaseline records the largest error, in float32 units in the last
place, of MLX's nine polynomial functions on the scalar CPU code, measured
on arm64 without Highway: exp, sin and cos (also over a wider range), erf,
erfInverse, sigmoid and logAddExp. PolynomialAccuracyTests allows a
Highway build 2 units more: the bound for Highway. The commit also adds
ReferenceMath, the references and checks the kernel tests share.
Affine weights from 32 rows on, the floating-point modes and the int8
path run at every Highway target the host runs, one target at a time;
below 32 rows, affine weights take code that is not dispatched.
Half-precision and non-contiguous activations run once, at the target
dispatch picks.
The kernels run at every Highway target the host runs, one target at a
time. float64, and the attention shapes the CPU kernel declines, go to
MLX's fallback, and the tests check that they do.
Where the result is determined, the tests compare it exactly, with +0
equal to -0 and any NaN matching any NaN; the functions MLX takes from
libm are compared with glibc's own. Elsewhere they check derived bounds:
MLX's polynomials in bf16 and fp16, float32 sums, softmax and logsumexp.
linux_swiftpm moves to swift:6.4-bookworm, the toolchain macOS builds use, and
runs scripts/run-cpu-kernel-tests.sh after VMLXLinuxTests: on x86-64 it tests
every Highway target the runner's CPU supports, on arm64 that none is built.
linux_cpu_kernels_sde runs the per-target tests under Intel SDE as Sapphire
Rapids, for the AVX-512 targets hosted runners may not have.
The fork now pins OpenBLAS to one thread whenever the CPU pool exists, a pool
of one thread included. shimAnswersAgree still expected a pin only above one
thread, and failed on x86-64 with MLX_CPU_THREADS=1, where OpenBLAS was pinned.
vmlx links OpenBLAS on Linux, so a Highway build reports it pinned at every
pool size.
KernelLock.run turns the int8 switch off before every test body, so
shimAnswersAgree's check of the switch could not fail. KernelLock now reads the
switch once, before its first reset, and shimAnswersAgree asserts that value.
The runner unsets MLX_CPU_QUANTIZED_INT8, so the value is the default. On
x86-64, with MLX_CPU_QUANTIZED_INT8=1 in the environment, the test failed in
both pool configurations before the runner unset it, and passed after.
withinSumBound took its reference and magnitude from MLX's own float64 matmul.
On builds with Highway kernels that is matmul<double> in gemms/cblas.cpp, which
carries a copy of the float32 path's row and batch splits, so a defect common
to both copies would have passed. hostMatmul computes a·b and |a|·|b| on the
host in Double, over any leading batch dimensions. withinQuantizedBound keeps
MLX's matmul: the kernel it judges is not cblas.
The runner exited 1 for a failed test, for a filter that matched no test (the
executable's 69) and for a stall (137 after the watcher's kill). It now exits
with the first failing configuration's status: 124 for a stall, 69 when no
test matched, and otherwise the executable's own, 1 for a failed test. It exits
2 when the probe's compiler ($CXX, by default g++) is missing, as for a usage
error, which --expect, --sde, --filter or --stall as the last argument now is,
where set -u stopped the script before. The default stall drops from 3600 s to
1200 s, so that a stall in each configuration still reports within the CI
job's 120 minutes. It picks the newest test executable, should a stale one
remain, and unsets VMLX_ULP_BASELINE, so that the polynomial bounds are judged.

On x86-64, with --skip-build: no match exited 69 (1 before), a test process
stopped with SIGSTOP 124 under --stall 15 (1 before), a wrong --expect 1 and
a missing CXX 2. --expect as the last argument exited 2 (1 before). With
VMLX_ULP_BASELINE=1 exported, the polynomial tests printed no measurement.
Without Highway kernels, five tests returned early and counted as passed: the
bf16 and fp16 cases of floatingPointModesWithinTheFloat32Bound,
halfPrecisionActivations, int8ActivationsOnlyWhenSwitchedOn,
halfPrecisionSumsAccumulateInFloat, and polynomialsInHalfPrecision's sigmoid
case. They now carry .disabled(if:) with the reason, as VMLXLinuxTests does. The
floating-point modes' bf16 and fp16 activations move to
floatingPointModesWithHalfPrecisionActivations, and sigmoid to
sigmoidInHalfPrecision, so the target has 46 tests instead of 44. On x86-64 all
46 ran and passed, with both pool configurations.
With VMLX_ULP_BASELINE=1 the polynomial tests printed their measurements and
passed, judging nothing. Each measurement now also records an issue, so that a
measuring run cannot pass for a judged one; the runner unsets the variable. On
x86-64 the measuring run printed all nine measurements and failed with 9
issues, where it had passed before.
polynomialsInHalfPrecision bounded each function by k·max(ulp32, 2^-33) plus
one unit of the dtype, where k is the scalar code's measured error in units.
erf's k was measured with units of at least 2^-24, since erf near 0 is accurate
only in absolute terms, so with a 2^-33 floor its bound near 0 claimed more
than the measurement supports. Each function now takes the floor its baseline
was measured with (polynomialFloor), which changes only erf's. On x86-64,
polynomialsInHalfPrecision passed with both pool configurations.
withinBound decided with MLX's own float64 operations and all(), which are
code under test, and failed with a bare false. The GEMM and quantized matmul
bounds now build a float64 reference per element (sumBound, matmulBound,
quantizedBound), and expectWithin judges it on the host through doubles(), as
the other kernel tests do, naming the element furthest outside its bound.
outsideBound serves the int8 test, which needs the float32 bound to fail. The
bounds themselves do not change, and each reference is computed once, outside
the per-target loop. On x86-64 all 46 tests passed with both pool
configurations.
- The dispatch counters indexed highway_info's table without a range check, so
  a family outside vmlx_cpu_family, VMLX_CPU_FAMILY_COUNT included, read past
  it. Such a family now has no counts. One static_assert per enumerator checks
  that vmlx_cpu_family lists highway_info::Family in order; the one on the
  count checked only the length.
- vmlx_cpu_highway_set_targets_for_test accepted any mask, so a test could
  restrict dispatch to a target whose instructions the CPU lacks. It now
  returns false for a mask with such a target, and leaves dispatch
  unrestricted.
- Without Highway kernels, the shim reported the int8 switch as a constant
  false, while Linux compiles precision.cpp in every build, which reads
  MLX_CPU_QUANTIZED_INT8. On Linux the shim now reports and sets the core's own
  switch; macOS keeps the constant.

On x86-64 all 47 tests passed with both pool configurations, the new
shimRefusesWhatItCannotServe among them. With the range and mask checks
removed it failed in both: the shim accepted the 57 target bits this CPU lacks,
and in the default pool the family past the table read 1. Swapping two
enumerators in the header stopped the build at their static_asserts. On arm64
without Highway kernels, a program linking the shim with precision.cpp read,
under MLX_CPU_QUANTIZED_INT8=1, 0 through the shim and 1 in the core before
this change, and 1 in both after.
- In a manifest, #if arch(x86_64) tests the host, so cross-compiling from
  x86-64 would define MLX_USE_HIGHWAY_KERNELS for a target that mlx's own CMake
  build refuses Highway for. highway-runtime/hwy_targets.cpp now stops such a
  build with an #error. Compiled for aarch64 with the macro defined, the
  wrapper passed before this change and fails with the error after.
- vmlx's CMake existence check now covers Source/Cmlx/highway, as it covers
  mlx and mlx-c. In a tree without it, configure failed in mlx's
  add_subdirectory ("not an existing directory") before, and stops at the
  check, with its message, after.
- NormTests: the RMSNorm bound, (n + 8)·ε·|y|, allows 4(n + 8)/(n + 10)
  times the analysis, 3.3 at n = 1, not 4.
- GemmTests: the batched bound's c reads 64 + 2, as K + 2 does elsewhere.
- The Highway runtime wrappers say why these six sources: the hwy library's,
  less its benchmarking and profiling tools. They no longer name the platform.
- Package.swift: the test target's comment says what the references are:
  Double on the host, except the quantized matmul bound's float64 matmul and
  the libm functions, compared with glibc's own.
- KernelLock.run runs each test body with the CPU as the default device.

On x86-64 all 47 kernel tests passed with both pool configurations.
expectIdentical compared values with ==, under which +0 equals -0, so a
kernel that lost the sign of a zero passed. It now compares bit patterns:
+0 and -0 differ, and any NaN still matches any NaN. doubles converts
exactly, so a zero keeps its sign on the way.

On x86-64 the strict comparison failed correctlyRoundedUnary for ceil and
round in float32, float64, bfloat16 and float16, 8 issues in each pool
configuration:

  ceil float32: 23 of 4099 differ; element 445 is 0.0, not -0.0
  round float32: 10 of 4099 differ; element 737 is 0.0, not -0.0

Below SSE4, Highway's Ceil and Round return +0 there, and the fork's
facade runs at SSE2. The fork's e5491a335 ("Keep the sign of zero in the
Highway facade's floor, ceil and rint") copies the input's sign onto their
results. With it, all 47 kernel tests passed with both pool configurations
on x86-64 and on arm64.
gatherQuantizedMM, the matmul of mixture-of-experts layers, had no kernel
test. #3019 multiplies consecutive routes to one expert, over rows that
follow on in x, as one matrix, which takes its dequantizing path from 32
rows on. For mxfp4, nvfp4 and mxfp8 with one row a route, and more routes
than pool threads, it hands whole routes to the pool's threads.

The routes, sorted by expert, make matrices of 40, 5, 33 and 36 routes, a
run that stays with an expert over rows that do not follow on, and single
routes, until they outnumber the threads twice over. On exact data the
affine experts must match float64 bit for bit; the floating-point experts
stay within the float32 bound of the other quantized tests, with bf16 and
fp16 activations on Highway builds only. All run with one and two rows a
route, at every Highway target the host runs. The references dequantize
the packed codes on the host, in Double (hostDequantized), without MLX's
dequantize. The counters show the target each family ran, not whether the
pool took the routes.

On x86-64 the new tests passed with both pool configurations, and on arm64,
without Highway kernels, those that run there did. With one route of the
run of 40 given the wrong expert in the reference, the affine test failed
at every target, with 20 issues in each pool configuration:

  bits 4, 1 rows a route, on AVX2: 128 of 17408 differ; element 2560 is 19.5, not -25.875

With the last route misrouted, the floating-point tests failed in all 90
cases:

  mxfp4 float32, 1 rows a route, on AVX2: element 8651 is -44.83448791503906, float64 gives 37.50072002701927, 4468.285207098716 times its bound
Two paths of the CPU quantized matmul had no kernel test. Untransposed
weights, x·w, take MLX's scalar _qmm and fp_qmm in every build. 2-bit
weights from 32 rows on take #3019's dequantize-then-SGEMM path with its
scalar dequantizer, since the dispatched one serves 4 and 8 bits.

On exact data, untransposed affine weights of 2, 4 and 8 bits, and 2-bit
weights in groups of 32, 64 and 128 at 32, 33 and 64 rows, must match
float64 bit for bit. Untransposed mxfp4, nvfp4 and mxfp8 stay within the
float32 bound. The references dequantize on the host, in Double. With
Highway kernels the tests also check that no dispatched affine kernel ran,
and that the floating-point modes counted their fallback.

On x86-64 and arm64 the tests passed with both pool configurations. With
one weight 1/8 off in the reference, the untransposed and the 2-bit tests
failed in all 9 cases each:

  bits 2 rows 1: 1 of 512 differ; element 70 is 197.0, not 198.0
  group 32 rows 32: 30 of 4096 differ; element 70 is 99.5, not 100.5

With a column of weights of the wrong sign, the floating-point test failed
in all 9 cases:

  mxfp4 rows 1: element 0 is 12.141921043395996, float64 gives -12.141925058647757, 4718.906759953502 times its bound
#3019 splits gather, gather_axis and copy_general_general across the pool
from 262144 elements on, and no kernel test reached those paths. An
embedding lookup of 512 tokens into a 768-wide table is 393216 elements.

IndexingTests runs each past that threshold, with every value its own
position in its source, and compares the result bit for bit with positions
computed on the host: take along axis 0, with ids as a vector, a batch and
a transposed batch, some negative, from rows that are not contiguous, and
of single elements; takeAlong along the last axis and along the middle
axis, of an array and of a transposed view; and the general copies of a
5-D view whose axes collapse nowhere, into a contiguous array and into
strided slices of a concatenation. They run with MLX_CPU_THREADS=1, on one
thread, and with the default pool, where #3019 splits them; no counter
records the split.

On x86-64 and arm64 the tests passed with both pool configurations. With
the last index shifted by one in each reference, the gathers failed in all
5 and 3 cases:

  ids as a vector: 768 of 393216 differ; element 392448 is 493824.0, not 494592.0
  the last axis: 1 of 524288 differ; element 524287 is 523671.0, not 523672.0

With two positions of the copy's last outer iteration swapped in the
reference, the copies failed in both cases:

  a strided source: 2 of 327680 differ; element 327678 is 327678.0, not 327679.0
Without F16C, at SSE4, SSSE3 and SSE2, Highway converts float16 to float32
in software, and its conversion takes exponent 31 for a finite number:
inf's bit pattern decodes to 65536. highway_utils.h's promote_f16 gives
inf and NaN there for the dispatched kernels' loads, but only the SIMD
facade's copy of that fix had a test.

RMSNorm and LayerNorm now take float16 rows with an infinity or a NaN,
made from their bits, at every Highway target the host runs. RMSNorm must
give NaN at an infinity and elsewhere a zero with the sign of x·w, as
float64 does, and NaN throughout a row with a NaN; LayerNorm NaN
throughout. The outputs are decoded from their bits on the host and
compared bit for bit, any NaN matching any NaN.

On x86-64 the test passed at every target the host runs (AVX2, SSE4,
SSSE3, SSE2 and EMU128) with both pool configurations, and on arm64. With
the reference reading exponent 31 as the software conversion did, it
failed at each of them:

  float16 rms_norm on SSE2: 480 of 480 differ; element 0 is -0.0, not -9.128146886085878e-05
  float16 layer_norm on SSE4: 480 of 480 differ; element 0 is nan, not 0.0018856052750296504
Comments, two failure messages and three manifest comments cited plan
tasks, spec sections and a tracker card that do not ship with this package.
Each now gives its reason: where a baseline was measured and over which
range, which build lacks a path and why, what switches the int8 path on.
The ULP baseline says how to measure it again, with VMLX_ULP_BASELINE=1 on
a build without Highway kernels, and no longer names an mlx commit that is
not on osaurus-ai/mlx. The bf16 reduction comment gives the sums logged
with partials rounded to bf16, 0 with 8 threads and -8192 with 9 or 12, in
place of a claim about every pool of 2 to 64 threads. Only comments and
message strings change.

The kernel tests passed with both pool configurations, 57 tests in 12
suites, on x86-64 and on arm64.
Source/Cmlx/mlx moves from e5491a335 to 1cf161ade (hwy/adopt-x86), which
claims a pool slot with one compare-and-swap, so that a stalled worker
cannot claim a slot of the next call, and reads the pool's generation
before a worker announces itself ready, so that no worker misses the
first call.
Source/Cmlx/mlx moves from 1cf161ade to 0203287b5 (hwy/adopt-x86), which
compares each lane with zero when a vector becomes Simd<bool>, so that
from_fp8 keeps the sign of negative values on every target. The Xcode
project leaves out the new tests/highway_simd_tests.cpp, as it leaves
out the fork's other tests.
eval_impl's back-pressure check reads the scheduler's and the
allocator's counters without their locks. ThreadSanitizer reports both
in every threaded run; neither is in the CPU pool or Highway's kernels,
which these runs exist to check.
Package.swift defines the Highway macros on Linux arm64 as on x86-64,
with HWY_DISABLED_TARGETS=HWY_ALL_SVE: the dispatched kernels are not
vector-length agnostic yet. The define is set on both architectures,
since a manifest's #if arch tests the host, not the target, and it does
nothing on x86-64. On arm64 the kernels then compile for
NEON_WITHOUT_AES, NEON and NEON_BF16. The Highway runtime's architecture
check accepts arm64, and the kernel-test runner expects those three
targets and EMU128 where the CPU supports them.
VMLX_NO_HIGHWAY=1 builds Linux without the Highway kernels and the CPU
pool: MLX's scalar CPU code, which vmlx ran on arm64 until now and
ULPBaseline is measured on. The kernel-test runner expects no Highway
target in such a build and gives it a scratch path of its own, and the
baseline's measuring mode refuses a Highway build. On arm64 the runner
names the CPU and the features that decide Highway's NEON targets, since
ARM's /proc/cpuinfo names no model. NormTests' float16 note now lists
EMU128 among the targets that convert in software, on x86-64 too, and
says arm64's NEON targets convert with FCVTL.
fromFP8 of every E4M3 code, and dequantize for mxfp4, nvfp4 and mxfp8,
against decodes written from the formats; maximum and minimum on +0, -0
and NaN, and max and min reductions with a NaN at every position.

The mxfp8 matmul test decoded its oracle with MLX's dequantize, so a
from_fp8 defect could fail it while the kernel was right, as the
Simd<bool> conversion that osaurus-ai/mlx#10 fixes did with Highway on
arm64, or pass a kernel that shared the defect. It now decodes on the
host, as the gather test does.

ExactBinary's maximum and minimum now give NaN for a NaN in either
operand. No verdict changes: the old form reached b's NaN through its
comparison, and expectIdentical treats any two NaNs as equal.
The aarch64 SwiftPM leg now builds the Highway kernels and expects the
NEON targets on GitHub's Neoverse N2 runner. A third leg builds without
Highway (VMLX_NO_HIGHWAY=1), the coverage of MLX's scalar CPU code the
aarch64 leg used to give. An advisory job runs the threaded kernel tests
under ThreadSanitizer on arm64, through the kernel-test runner's new
--sanitize thread, which applies scripts/tsan-suppressions.txt, prints
any sanitizer summary and warns when no suppression matched; the job
uploads the logs when it fails.
Source/Cmlx/mlx moves from fdb6bbacf to 4c12dccc0: the Highway kernels
on Linux arm64 with the SVE family disabled, maximum and minimum with
the scalar rule for +0 and -0 on Highway targets, and mask storage of
the size Highway's contract asks for.
@rcfa
rcfa changed the base branch from main to ci-base-t167p2 October 1, 2026 03:42
@rcfa rcfa closed this Oct 1, 2026
@rcfa rcfa reopened this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant