Repository navigation
Conversation
scripts/check-macos-identity.py rebuilds two commits' trees from git and compares the sources SwiftPM compiles for Cmlx on macOS, each preprocessed under the manifest's macOS settings, and the files the Xcode project compiles from its mlx and mlx-c folders. vmlx's CMake build is outside it.
Move Source/Cmlx/mlx to the osaurus-ai/mlx pull request that adopts ml-explore/mlx#3019's Highway kernels, and compile them on Linux x86-64: MLX_USE_HIGHWAY_KERNELS, the Highway runtime through wrappers in highway-runtime/, and the submodule's headers. macOS excludes every new source, in SwiftPM and in the Xcode project, so both compile what they did; vmlx's CMake build compiles mlx's precision.cpp, the int8 switch, everywhere, macOS included. Linux arm64 compiles the new sources as empty files, all but the int8 switch, until the kernels are tested there. VMLX_HWY_ALL_TARGETS=1 adds HWY_COMPILE_ALL_ATTAINABLE for the kernel tests. vmlx's CMake build takes Highway from the submodule.
Swift reads and sets them through it: the Highway targets, the int8 quantized-matmul switch, the thread pool and the per-family dispatch counters. Builds without Highway kernels answer as such a build does: no targets, int8 off, one thread, no counts.
The CPU backend's kernel tests run on Linux under one lock, since the Highway target mask, the int8 switch and the dispatch counters are process-wide. scripts/run-cpu-kernel-tests.sh builds the all-targets test build and runs them twice, with MLX_CPU_THREADS=1 and with the default pool. The targets a host must execute come from Highway's own detection of its CPU, intersected with the targets Highway attains there, never from what the build contains.
The runner's failure excerpt printed swift-testing's summary lines but not the values it prints under each recorded issue, so a failure showed that an expectation failed, not with what values. Print each recorded issue with its values, and any stall line.
ULPBaseline records the largest error, in float32 units in the last place, of MLX's nine polynomial functions on the scalar CPU code, measured on arm64 without Highway: exp, sin and cos (also over a wider range), erf, erfInverse, sigmoid and logAddExp. PolynomialAccuracyTests allows a Highway build 2 units more: the bound for Highway. The commit also adds ReferenceMath, the references and checks the kernel tests share.
Affine weights from 32 rows on, the floating-point modes and the int8 path run at every Highway target the host runs, one target at a time; below 32 rows, affine weights take code that is not dispatched. Half-precision and non-contiguous activations run once, at the target dispatch picks.
The kernels run at every Highway target the host runs, one target at a time. float64, and the attention shapes the CPU kernel declines, go to MLX's fallback, and the tests check that they do.
Where the result is determined, the tests compare it exactly, with +0 equal to -0 and any NaN matching any NaN; the functions MLX takes from libm are compared with glibc's own. Elsewhere they check derived bounds: MLX's polynomials in bf16 and fp16, float32 sums, softmax and logsumexp.
linux_swiftpm moves to swift:6.4-bookworm, the toolchain macOS builds use, and runs scripts/run-cpu-kernel-tests.sh after VMLXLinuxTests: on x86-64 it tests every Highway target the runner's CPU supports, on arm64 that none is built. linux_cpu_kernels_sde runs the per-target tests under Intel SDE as Sapphire Rapids, for the AVX-512 targets hosted runners may not have.
The fork now pins OpenBLAS to one thread whenever the CPU pool exists, a pool of one thread included. shimAnswersAgree still expected a pin only above one thread, and failed on x86-64 with MLX_CPU_THREADS=1, where OpenBLAS was pinned. vmlx links OpenBLAS on Linux, so a Highway build reports it pinned at every pool size.
KernelLock.run turns the int8 switch off before every test body, so shimAnswersAgree's check of the switch could not fail. KernelLock now reads the switch once, before its first reset, and shimAnswersAgree asserts that value. The runner unsets MLX_CPU_QUANTIZED_INT8, so the value is the default. On x86-64, with MLX_CPU_QUANTIZED_INT8=1 in the environment, the test failed in both pool configurations before the runner unset it, and passed after.
withinSumBound took its reference and magnitude from MLX's own float64 matmul. On builds with Highway kernels that is matmul<double> in gemms/cblas.cpp, which carries a copy of the float32 path's row and batch splits, so a defect common to both copies would have passed. hostMatmul computes a·b and |a|·|b| on the host in Double, over any leading batch dimensions. withinQuantizedBound keeps MLX's matmul: the kernel it judges is not cblas.
The runner exited 1 for a failed test, for a filter that matched no test (the executable's 69) and for a stall (137 after the watcher's kill). It now exits with the first failing configuration's status: 124 for a stall, 69 when no test matched, and otherwise the executable's own, 1 for a failed test. It exits 2 when the probe's compiler ($CXX, by default g++) is missing, as for a usage error, which --expect, --sde, --filter or --stall as the last argument now is, where set -u stopped the script before. The default stall drops from 3600 s to 1200 s, so that a stall in each configuration still reports within the CI job's 120 minutes. It picks the newest test executable, should a stale one remain, and unsets VMLX_ULP_BASELINE, so that the polynomial bounds are judged. On x86-64, with --skip-build: no match exited 69 (1 before), a test process stopped with SIGSTOP 124 under --stall 15 (1 before), a wrong --expect 1 and a missing CXX 2. --expect as the last argument exited 2 (1 before). With VMLX_ULP_BASELINE=1 exported, the polynomial tests printed no measurement.
Without Highway kernels, five tests returned early and counted as passed: the bf16 and fp16 cases of floatingPointModesWithinTheFloat32Bound, halfPrecisionActivations, int8ActivationsOnlyWhenSwitchedOn, halfPrecisionSumsAccumulateInFloat, and polynomialsInHalfPrecision's sigmoid case. They now carry .disabled(if:) with the reason, as VMLXLinuxTests does. The floating-point modes' bf16 and fp16 activations move to floatingPointModesWithHalfPrecisionActivations, and sigmoid to sigmoidInHalfPrecision, so the target has 46 tests instead of 44. On x86-64 all 46 ran and passed, with both pool configurations.
With VMLX_ULP_BASELINE=1 the polynomial tests printed their measurements and passed, judging nothing. Each measurement now also records an issue, so that a measuring run cannot pass for a judged one; the runner unsets the variable. On x86-64 the measuring run printed all nine measurements and failed with 9 issues, where it had passed before.
polynomialsInHalfPrecision bounded each function by k·max(ulp32, 2^-33) plus one unit of the dtype, where k is the scalar code's measured error in units. erf's k was measured with units of at least 2^-24, since erf near 0 is accurate only in absolute terms, so with a 2^-33 floor its bound near 0 claimed more than the measurement supports. Each function now takes the floor its baseline was measured with (polynomialFloor), which changes only erf's. On x86-64, polynomialsInHalfPrecision passed with both pool configurations.
withinBound decided with MLX's own float64 operations and all(), which are code under test, and failed with a bare false. The GEMM and quantized matmul bounds now build a float64 reference per element (sumBound, matmulBound, quantizedBound), and expectWithin judges it on the host through doubles(), as the other kernel tests do, naming the element furthest outside its bound. outsideBound serves the int8 test, which needs the float32 bound to fail. The bounds themselves do not change, and each reference is computed once, outside the per-target loop. On x86-64 all 46 tests passed with both pool configurations.
- The dispatch counters indexed highway_info's table without a range check, so a family outside vmlx_cpu_family, VMLX_CPU_FAMILY_COUNT included, read past it. Such a family now has no counts. One static_assert per enumerator checks that vmlx_cpu_family lists highway_info::Family in order; the one on the count checked only the length. - vmlx_cpu_highway_set_targets_for_test accepted any mask, so a test could restrict dispatch to a target whose instructions the CPU lacks. It now returns false for a mask with such a target, and leaves dispatch unrestricted. - Without Highway kernels, the shim reported the int8 switch as a constant false, while Linux compiles precision.cpp in every build, which reads MLX_CPU_QUANTIZED_INT8. On Linux the shim now reports and sets the core's own switch; macOS keeps the constant. On x86-64 all 47 tests passed with both pool configurations, the new shimRefusesWhatItCannotServe among them. With the range and mask checks removed it failed in both: the shim accepted the 57 target bits this CPU lacks, and in the default pool the family past the table read 1. Swapping two enumerators in the header stopped the build at their static_asserts. On arm64 without Highway kernels, a program linking the shim with precision.cpp read, under MLX_CPU_QUANTIZED_INT8=1, 0 through the shim and 1 in the core before this change, and 1 in both after.
- In a manifest, #if arch(x86_64) tests the host, so cross-compiling from
x86-64 would define MLX_USE_HIGHWAY_KERNELS for a target that mlx's own CMake
build refuses Highway for. highway-runtime/hwy_targets.cpp now stops such a
build with an #error. Compiled for aarch64 with the macro defined, the
wrapper passed before this change and fails with the error after.
- vmlx's CMake existence check now covers Source/Cmlx/highway, as it covers
mlx and mlx-c. In a tree without it, configure failed in mlx's
add_subdirectory ("not an existing directory") before, and stops at the
check, with its message, after.
- NormTests: the RMSNorm bound, (n + 8)·ε·|y|, allows 4(n + 8)/(n + 10) times the analysis, 3.3 at n = 1, not 4. - GemmTests: the batched bound's c reads 64 + 2, as K + 2 does elsewhere. - The Highway runtime wrappers say why these six sources: the hwy library's, less its benchmarking and profiling tools. They no longer name the platform. - Package.swift: the test target's comment says what the references are: Double on the host, except the quantized matmul bound's float64 matmul and the libm functions, compared with glibc's own. - KernelLock.run runs each test body with the CPU as the default device. On x86-64 all 47 kernel tests passed with both pool configurations.
expectIdentical compared values with ==, under which +0 equals -0, so a
kernel that lost the sign of a zero passed. It now compares bit patterns:
+0 and -0 differ, and any NaN still matches any NaN. doubles converts
exactly, so a zero keeps its sign on the way.
On x86-64 the strict comparison failed correctlyRoundedUnary for ceil and
round in float32, float64, bfloat16 and float16, 8 issues in each pool
configuration:
ceil float32: 23 of 4099 differ; element 445 is 0.0, not -0.0
round float32: 10 of 4099 differ; element 737 is 0.0, not -0.0
Below SSE4, Highway's Ceil and Round return +0 there, and the fork's
facade runs at SSE2. The fork's e5491a335 ("Keep the sign of zero in the
Highway facade's floor, ceil and rint") copies the input's sign onto their
results. With it, all 47 kernel tests passed with both pool configurations
on x86-64 and on arm64.
gatherQuantizedMM, the matmul of mixture-of-experts layers, had no kernel test. #3019 multiplies consecutive routes to one expert, over rows that follow on in x, as one matrix, which takes its dequantizing path from 32 rows on. For mxfp4, nvfp4 and mxfp8 with one row a route, and more routes than pool threads, it hands whole routes to the pool's threads. The routes, sorted by expert, make matrices of 40, 5, 33 and 36 routes, a run that stays with an expert over rows that do not follow on, and single routes, until they outnumber the threads twice over. On exact data the affine experts must match float64 bit for bit; the floating-point experts stay within the float32 bound of the other quantized tests, with bf16 and fp16 activations on Highway builds only. All run with one and two rows a route, at every Highway target the host runs. The references dequantize the packed codes on the host, in Double (hostDequantized), without MLX's dequantize. The counters show the target each family ran, not whether the pool took the routes. On x86-64 the new tests passed with both pool configurations, and on arm64, without Highway kernels, those that run there did. With one route of the run of 40 given the wrong expert in the reference, the affine test failed at every target, with 20 issues in each pool configuration: bits 4, 1 rows a route, on AVX2: 128 of 17408 differ; element 2560 is 19.5, not -25.875 With the last route misrouted, the floating-point tests failed in all 90 cases: mxfp4 float32, 1 rows a route, on AVX2: element 8651 is -44.83448791503906, float64 gives 37.50072002701927, 4468.285207098716 times its bound
Two paths of the CPU quantized matmul had no kernel test. Untransposed weights, x·w, take MLX's scalar _qmm and fp_qmm in every build. 2-bit weights from 32 rows on take #3019's dequantize-then-SGEMM path with its scalar dequantizer, since the dispatched one serves 4 and 8 bits. On exact data, untransposed affine weights of 2, 4 and 8 bits, and 2-bit weights in groups of 32, 64 and 128 at 32, 33 and 64 rows, must match float64 bit for bit. Untransposed mxfp4, nvfp4 and mxfp8 stay within the float32 bound. The references dequantize on the host, in Double. With Highway kernels the tests also check that no dispatched affine kernel ran, and that the floating-point modes counted their fallback. On x86-64 and arm64 the tests passed with both pool configurations. With one weight 1/8 off in the reference, the untransposed and the 2-bit tests failed in all 9 cases each: bits 2 rows 1: 1 of 512 differ; element 70 is 197.0, not 198.0 group 32 rows 32: 30 of 4096 differ; element 70 is 99.5, not 100.5 With a column of weights of the wrong sign, the floating-point test failed in all 9 cases: mxfp4 rows 1: element 0 is 12.141921043395996, float64 gives -12.141925058647757, 4718.906759953502 times its bound
#3019 splits gather, gather_axis and copy_general_general across the pool from 262144 elements on, and no kernel test reached those paths. An embedding lookup of 512 tokens into a 768-wide table is 393216 elements. IndexingTests runs each past that threshold, with every value its own position in its source, and compares the result bit for bit with positions computed on the host: take along axis 0, with ids as a vector, a batch and a transposed batch, some negative, from rows that are not contiguous, and of single elements; takeAlong along the last axis and along the middle axis, of an array and of a transposed view; and the general copies of a 5-D view whose axes collapse nowhere, into a contiguous array and into strided slices of a concatenation. They run with MLX_CPU_THREADS=1, on one thread, and with the default pool, where #3019 splits them; no counter records the split. On x86-64 and arm64 the tests passed with both pool configurations. With the last index shifted by one in each reference, the gathers failed in all 5 and 3 cases: ids as a vector: 768 of 393216 differ; element 392448 is 493824.0, not 494592.0 the last axis: 1 of 524288 differ; element 524287 is 523671.0, not 523672.0 With two positions of the copy's last outer iteration swapped in the reference, the copies failed in both cases: a strided source: 2 of 327680 differ; element 327678 is 327678.0, not 327679.0
Without F16C, at SSE4, SSSE3 and SSE2, Highway converts float16 to float32 in software, and its conversion takes exponent 31 for a finite number: inf's bit pattern decodes to 65536. highway_utils.h's promote_f16 gives inf and NaN there for the dispatched kernels' loads, but only the SIMD facade's copy of that fix had a test. RMSNorm and LayerNorm now take float16 rows with an infinity or a NaN, made from their bits, at every Highway target the host runs. RMSNorm must give NaN at an infinity and elsewhere a zero with the sign of x·w, as float64 does, and NaN throughout a row with a NaN; LayerNorm NaN throughout. The outputs are decoded from their bits on the host and compared bit for bit, any NaN matching any NaN. On x86-64 the test passed at every target the host runs (AVX2, SSE4, SSSE3, SSE2 and EMU128) with both pool configurations, and on arm64. With the reference reading exponent 31 as the software conversion did, it failed at each of them: float16 rms_norm on SSE2: 480 of 480 differ; element 0 is -0.0, not -9.128146886085878e-05 float16 layer_norm on SSE4: 480 of 480 differ; element 0 is nan, not 0.0018856052750296504
Comments, two failure messages and three manifest comments cited plan tasks, spec sections and a tracker card that do not ship with this package. Each now gives its reason: where a baseline was measured and over which range, which build lacks a path and why, what switches the int8 path on. The ULP baseline says how to measure it again, with VMLX_ULP_BASELINE=1 on a build without Highway kernels, and no longer names an mlx commit that is not on osaurus-ai/mlx. The bf16 reduction comment gives the sums logged with partials rounded to bf16, 0 with 8 threads and -8192 with 9 or 12, in place of a claim about every pool of 2 to 64 threads. Only comments and message strings change. The kernel tests passed with both pool configurations, 57 tests in 12 suites, on x86-64 and on arm64.
Source/Cmlx/mlx moves from e5491a335 to 1cf161ade (hwy/adopt-x86), which claims a pool slot with one compare-and-swap, so that a stalled worker cannot claim a slot of the next call, and reads the pool's generation before a worker announces itself ready, so that no worker misses the first call.
Source/Cmlx/mlx moves from 1cf161ade to 0203287b5 (hwy/adopt-x86), which compares each lane with zero when a vector becomes Simd<bool>, so that from_fp8 keeps the sign of negative values on every target. The Xcode project leaves out the new tests/highway_simd_tests.cpp, as it leaves out the fork's other tests.
eval_impl's back-pressure check reads the scheduler's and the allocator's counters without their locks. ThreadSanitizer reports both in every threaded run; neither is in the CPU pool or Highway's kernels, which these runs exist to check.
Package.swift defines the Highway macros on Linux arm64 as on x86-64, with HWY_DISABLED_TARGETS=HWY_ALL_SVE: the dispatched kernels are not vector-length agnostic yet. The define is set on both architectures, since a manifest's #if arch tests the host, not the target, and it does nothing on x86-64. On arm64 the kernels then compile for NEON_WITHOUT_AES, NEON and NEON_BF16. The Highway runtime's architecture check accepts arm64, and the kernel-test runner expects those three targets and EMU128 where the CPU supports them.
VMLX_NO_HIGHWAY=1 builds Linux without the Highway kernels and the CPU pool: MLX's scalar CPU code, which vmlx ran on arm64 until now and ULPBaseline is measured on. The kernel-test runner expects no Highway target in such a build and gives it a scratch path of its own, and the baseline's measuring mode refuses a Highway build. On arm64 the runner names the CPU and the features that decide Highway's NEON targets, since ARM's /proc/cpuinfo names no model. NormTests' float16 note now lists EMU128 among the targets that convert in software, on x86-64 too, and says arm64's NEON targets convert with FCVTL.
fromFP8 of every E4M3 code, and dequantize for mxfp4, nvfp4 and mxfp8, against decodes written from the formats; maximum and minimum on +0, -0 and NaN, and max and min reductions with a NaN at every position. The mxfp8 matmul test decoded its oracle with MLX's dequantize, so a from_fp8 defect could fail it while the kernel was right, as the Simd<bool> conversion that osaurus-ai/mlx#10 fixes did with Highway on arm64, or pass a kernel that shared the defect. It now decodes on the host, as the gather test does. ExactBinary's maximum and minimum now give NaN for a NaN in either operand. No verdict changes: the old form reached b's NaN through its comparison, and expectIdentical treats any two NaNs as equal.
The aarch64 SwiftPM leg now builds the Highway kernels and expects the NEON targets on GitHub's Neoverse N2 runner. A third leg builds without Highway (VMLX_NO_HIGHWAY=1), the coverage of MLX's scalar CPU code the aarch64 leg used to give. An advisory job runs the threaded kernel tests under ThreadSanitizer on arm64, through the kernel-test runner's new --sanitize thread, which applies scripts/tsan-suppressions.txt, prints any sanitizer summary and warns when no suppression matched; the job uploads the logs when it fails.
Source/Cmlx/mlx moves from fdb6bbacf to 4c12dccc0: the Highway kernels on Linux arm64 with the SVE family disabled, maximum and minimum with the scalar rule for +0 and -0 on Highway targets, and mask storage of the size Highway's contract asks for.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Runs vmlx-swift's CI in this fork for plan 2 of T-167 phase 2: Highway on Linux arm64, a Linux leg without Highway, and an advisory ThreadSanitizer job. Not for merge; the upstream PR follows osaurus-ai/mlx#18.
🤖 Generated with Claude Code