Skip to content

arm64 SME2: faster SGEMM/DGEMM, and SYMM/SYRK/SYR2K/TRMM/TRSM on the SME GEMM kernel - #6072

Closed
tesch1 wants to merge 2 commits into
OpenMathLib:developfrom
tesch1:sme2-gemm
Closed

tesch1 wants to merge 2 commits into
OpenMathLib:developfrom
tesch1:sme2-gemm

Conversation

@tesch1

@tesch1 tesch1 commented Sep 29, 2026 •

Copy link
Copy Markdown

Summary

Closes #6073.

On SME2 hardware with a 512-bit streaming vector length (checked at run time), sme_sgemm_kernel and
sme_dgemm_kernel run a blocked SME2 GEMM after C. Deng, W. Yang, J. Fang, D. Dong, Demystifying ARM SME to
Optimize General Matrix Multiplications
(arXiv:2512.21473), as reproduced and measured in MTGEMM-A
(https://github.com/tesch1/mtgemm-a). The kernels from #5971 stay as the fallback for other SME hardware.
SYMM, SYRK, SYR2K, TRMM and TRSM, which ran on the NEON kernels of the level-3 driver, now use the SME GEMM kernel.

Apple M4 Pro (GFLOPS; GEMM rows are geometric means over the shape set; "default" is each library's own thread
count; all measured on the same machine with the same harness; row-major / column-major where two values are
given):

Accelerate develop (63d7f22) this PR MTGEMM-A
SGEMM, the paper's 24 workloads, 1 thread 1127 / 1185 707 / 766 1389 / 1391 1404 / 1408
SGEMM, same, default threads 2364 / 2445 702 / 766 2709 / 2731 2733 / 2770
SGEMM, squares 512-4096, 1 thread 1691 / 1624 1129 / 1342 1747 / 1721 1748 / 1727
SGEMM, squares 4-384, 1 thread 247 / 254 202 / 70 210 / 225 179 / 167
SGEMM, thin (M or N 1-64, rest 4096), 1 thread 198 / 197 118 / 49 268 / 263 273 / 270
DGEMM, the paper's 24 workloads, 1 thread 321 / 344 275 / 283 423 / 406 426 / 409
SSYMM / SSYRK / SSYR2K / STRMM / STRSM, n = 1024, 1 thread 1559 / 1758 / 1384 / 1537 / 791 115 / 108 / 110 / 112 / 99 1318 / 1317 / 1329 / 1107 / 584 -
same, default threads 2732 / 2845 / 2811 / 2643 / 1589 376 / 327 / 224 / 203 / 171 2350 / 2003 / 1966 / 1648 / 742 -
DSYMM / DSYRK / DSYR2K / DTRMM / DTRSM, n = 1024, 1 thread 410 / 403 / 391 / 433 / 265 58 / 55 / 55 / 57 / 53 398 / 387 / 374 / 358 / 257 -

Not slower than develop anywhere measured: GEMM on 348 shape/configuration pairs and, per call, 54 square
sizes from 4^3 to 256^3 in both orders and precisions; SYMM, SYRK, SYR2K, TRMM and TRSM at n = 64-4096 on 1 thread
and default threads, and every variant at n = 1024. The only ratios below 1.0 (0.95x-0.99x per call) are sizes
where both libraries run the same code: row-major SGEMM below 64^3 stays on develop's SME1 direct kernel, fp64 16^3
on its SME kernel. That direct kernel mallocs a scratch copy on every call and its speed depends on where the
allocation lands; develop itself gives 82-137 GFLOPS at 24^3 and 230-400 at 32^3 from process to process.

Where this PR is still slower:

  • Than Accelerate, GEMM: small problems (fp32 4^3-8^3 at 0.25x-0.54x, fp32 24^3-96^3 at 0.78x-0.94x,
    fp64 4^3-16^3 at 0.50x-0.93x, fp64 24^3-96^3 at 0.60x-0.93x), shapes with M or N = 1 (0.36x-0.58x), fp64
    16 x 16 x 4096, 16 x 4096 x 4096 and 4096 x 16 x 4096 (0.74x-0.75x), and fp32 4096 x 24576 x 1536 at default
    threads (0.89x). Every other shape is at 0.95x or
    better, and the geometric means of the paper's workloads and large squares are 1.0x-1.3x Accelerate.
  • Than Accelerate, level 3: every routine up to n = 512 (SYMM, SYRK, SYR2K 0.28x-0.87x at 64-512, and SYRK
    at 64, which stays on develop's NEON path below the threshold, 0.04x-0.07x; TRMM, TRSM 0.16x-0.76x). At n = 1024,
    SYMM, SYRK, SYR2K and TRMM are 0.72x-0.97x on one thread and 0.62x-1.20x at default threads; from n = 2048
    0.74x-1.13x (1.61x for fp32 SYRK at 2048 at default threads, where Accelerate drops). TRSM from n = 1024 is
    0.74x-1.14x on one thread and 0.47x-1.10x at default threads, where its base solves still run on one thread;
    over the variants at n = 1024, fp32 577-591 GFLOPS (triangle on the left) and 673-696 (right) against
    Accelerate's 791-879 and 840-949; fp64 253-258 and 273-275 against 263-290 and 207-252.
  • Than MTGEMM-A (the reference implementation of the design): 0.99x-1.00x on the paper's workloads and large
    squares; per call 0.93x-0.95x at its worst small sizes and faster at 185 of 216 (NEON and the older kernels
    take the smallest problems); 0.97x-0.98x on fp32 thin shapes (memory-bound; within 2% when timed in one
    process).

Full tables and raw data: https://github.com/tesch1/mtgemm-a (README, "What this means for OpenBLAS").

Commits

  1. arm64: SME2 sgemm/dgemm kernel for 512-bit streaming vector length
    • kernel/arm64/sme2_gemm_impl.h (new), included by sme_sgemm_kernel.c and sme_dgemm_kernel.c: blocking
      from the paper's analytical model, the ZA transposition of A in packing, online packing of B during the first
      micro-kernel pass, a 1 x NT-tile micro-kernel with multi-vector loads (plus half-width and edge kernels with
      predicates for every column tail wider than one vector), four-row ZA moves for C (also partial-width),
      core-side L2 prefetch, all four transpose cases (vectorised copy for column-contiguous A, alpha applied after
      the ZA transposition), and one thread per SME unit through exec_blas from 2^22 multiply-adds.
    • #pragma clang attribute enables SME2 for these functions only, so no build flags change; the path is
      compiled with Apple clang >= 17 or LLVM >= 18 (sme2_gemm_detect.h also has the run-time check).
    • Small problems, measured per call from 4^3 to 256^3 against develop and MTGEMM-A: compile-time variants of
      the ZA transposition; a per-thread packing buffer (freed at thread exit; the stack and blas_memory_alloc
      were slower); interface/gemm.c keeps problems below 20^3 multiply-adds (fp64 16^3) on NEON; the SME1 direct
      SGEMM keeps row-major problems below 64^3 (checked first, so small calls cost what they did); fp64 below 5000
      with M and N multiples of 16 keeps the existing SME kernel.
    • SME units: Apple has one per performance cluster. M5 Pro/Max have no efficiency tier
      (hw.perflevel1.name is "Performance") and have been reported to have two SME units for one Super cluster,
      so the count doubles there (not measured here).
    • M, N or K of 2^30 or more (INTERFACE64) stay on the existing kernels, since the kernel keeps them in int;
      if blas_memory_alloc fails, the two-thread split runs on one thread and the workspace comes from malloc.
    • cpuid_arm64.c: the M4 Pro (hw.cpufamily 0x17d5b93a) was not recognised, so builds without TARGET fell
      back to ARMV8 without SME; later Apple cores reporting FEAT_SME now map to VORTEXM4.
  2. arm64: SYMM, SYRK, SYR2K, TRMM and TRSM on the SME GEMM kernel
    • kernel/arm64/sme_level3.h (new), for real types: the symmetric/triangular dimension is split recursively,
      off-diagonal blocks are GEMM calls on SME_[SD]GEMM_KERNEL, diagonal blocks (<= 64) run as one GEMM on a
      full copy (SYMM, SYRK, SYR2K, TRMM), and TRSM solves blocks <= 32 by substitution in a local tile of 16 rows
      (right side) or 16 columns (left side, moved with NEON 4 x 4 / 2 x 2 transposes); the generic TRSM kernel
      runs at 10-16 GFLOPS on such blocks. Every side/uplo/trans/diag, from 2e5 multiply-adds.
    • interface/{symm,syrk,syr2k,trsm}.c each get only an ARCH_ARM64-guarded include and one guarded call to
      s3_symm_hook, s3_syrk_hook or s3_trxm_hook (8 lines per file). TRMM falls back to the level-3 driver
      if no buffer is available.
    • README: the VortexM4 line names the SME2 kernel and the level-3 routines on it; CONTRIBUTORS.md entry.

Why the level-3 routines are not in the level-3 driver

The driver runs SYMM/SYRK/SYR2K/TRMM/TRSM through their own pack and micro-kernel routines, which share
GEMM_UNROLL_M/N with the GEMM kernel of the target (the constraint that held up #5011). The SME2 kernel packs
inside its own blocking, so sharing unroll sizes with it would change every other kernel of VORTEXM4 and
ARMV9SME. Recursive blocking over the complete GEMM kernel needs no new pack routines, no gotoblas_t members and
no build-system changes; it is the usual way to put these routines on a fast GEMM. The code is in
kernel/arm64/sme_level3.h; the interface files only call it, after the argument checks and quick returns and
after develop's SME1 direct kernels. Those still run first where they apply: row-major fp32 SSYMM (left side),
SSYRK and STRMM (left side, non-unit) with packed leading dimensions, and row-major SGEMM below 64^3.

Relation to #5011

#5011 (WIP SME2 SGEMM through the level-3 driver) stopped at the point this PR works around: SYMM and TRMM
share the GEMM unroll sizes, so a new GEMM kernel shape breaks them. On cores with a 512-bit streaming vector
length (all Apple SME cores so far) this PR covers SGEMM and DGEMM and the level-3 routines with SYMM/TRMM
passing, so #5011 could be closed in its favour for those cores. #5011 also targets other vector lengths,
which this PR does not handle (they keep the #5971 kernels); that part would remain open.

Testing

Machine: Apple M4 Pro (8 performance + 4 efficiency cores; two SME units, one per performance cluster),
macOS 27 (Darwin 27.0.0), Apple clang 21.0.0.

Untested on M5, M5 Pro/Max, M6 and all other SME hardware. On those, the new paths run only if the core
reports SME2 and a 512-bit streaming vector length (checked at run time); otherwise develop's kernels run
unchanged. DYNAMIC_ARCH already selects vortexm4 for every Apple core with FEAT_SME, and getarch now maps
unknown Apple cores with FEAT_SME to VORTEXM4. The thresholds and the SME unit count were tuned or derived on the
M4 Pro only; the M5 Pro/Max unit count comes from a third-party report and is not measured here.
Results from other machines would be welcome.

Build configurations

configuration flags result
static, pthreads TARGET=VORTEXM4 USE_THREAD=1 USE_OPENMP=0 NUM_THREADS=12 NO_LAPACK=1 all checks pass
Homebrew-style DYNAMIC_ARCH=1 USE_OPENMP=1 NUM_THREADS=56 TARGET=VORTEX NO_SVE=1, Apple clang with -Xpreprocessor -fopenmp and libomp, as the Homebrew formula builds all checks pass; core vortexm4 selected at run time, new paths taken
autodetection no TARGET getarch now reports VORTEXM4 on the M4 Pro (was ARMV8)

OpenBLAS's own tests

Both macOS configurations above, built with make and gfortran-16 (NO_LAPACK=1), pass test (s/d/c/z blat1-3),
ctest (cblas s/d/c/z, both orders) and utest (125 + 1500 tests). Their default sizes (up to 31) stay below
the new level-3 paths, so sblat3 and dblat3 were also run with N = 0 1 2 3 7 31 33 63 65 (65 is the testers'
NMAX), which reaches them: all six routines pass in both precisions and both configurations.

GEMM checks (test_openblas_gemm.c)

12,448 calls of cblas_sgemm and cblas_dgemm, each compared element by element with a long-double reference;
tolerance 4 (k + 4) eps (eps = 1.2e-7 fp32, 2.3e-16 fp64), inputs uniform in [-1, 1]:

  • both orders, all four transpose combinations, alpha/beta in {(1, 0), (1, 1), (-0.5, 0.25), (2, -1)};
  • M, N, K from {1, 3, 7, 16, 17, 31, 32, 33, 48, 63, 64, 65, 97, 128, 130, 200}, which cross every tile,
    panel and tail boundary of the kernels (16/32/64 fp32, 8/16/64 fp64), plus (256, 300, 260), (513, 129, 700),
    (70, 1100, 300), (1030, 40, 1500), (600, 700, 1300), which exercise blocking and the two-thread split;
  • leading dimensions padded by 0-2;
  • 1 thread and 12 threads (the two-thread SME path above 2^22 multiply-adds).

In addition every call of the GEMM benchmark (348 shape and configuration pairs) is checked against MTGEMM-A:
340 are bit for bit identical (as is develop on all 348); the other eight, fp32 column-major 4^3-16^3 on one
thread and at default threads, go to the NEON kernels and differ by at most 8.4e-8 sqrt(K).

Small sizes, per call

bench_small times 54 square sizes from 4^3 to 256^3 per call (minimum over many short batches, three passes over
all sizes, since a pass can start on an efficiency core whose SME unit is several times slower), for develop, this
PR and MTGEMM-A, one thread, both orders and precisions. The tables are under "Performance details".

Level-3 checks (test_openblas_l3.c)

2,816 calls, each against a long-double reference, in fp32 and fp64:

routine variants calls
SYMM order x side x uplo 256
SYRK order x uplo x trans 256
SYR2K order x uplo x trans 256
TRMM order x side x uplo x trans x diag 1024
TRSM order x side x uplo x trans x diag (checked through the residual op(A) X - alpha B) 1024
  • (m, n) (or (n, k)) from {(150, 130), (257, 90), (90, 300), (64, 64), (65, 200), (200, 33), (513, 70),
    (128, 128)}: below, at and above the recursion blocks (32 for TRSM, 64 otherwise) and odd splits;
  • four alpha/beta pairs, leading dimensions padded by 0-2, 1 and 12 threads;
  • the unused triangle of A, the diagonal when it is unit, and the padding of A hold NaN, so any read of them
    fails the check; every element of B or C outside the result (other triangle, padding) must come back bit for
    bit; TRSM uses a diagonally dominant A.

Develop passes the same test; the tests are in https://github.com/tesch1/mtgemm-a (bench/).

Performance details

Harness: minimum of five trials of at least 50 ms each, after a 12 s wait (idle OpenBLAS pool threads spin for
2^28 timer ticks, about 11 s at Apple's 24 MHz, after start-up and slow the first calls by up to 25%); threads at
QoS user-interactive; runs start only below a one-minute load average of 4. Timings of the same build vary by
up to about 3% between sessions.

GEMM, per shape

Accelerate / develop / this PR / MTGEMM-A, GFLOPS; "row" is row-major C = A*B (beta 0), "col" column-major C += A*B (beta 1).

The paper's 24 workloads and squares 512-4096, fp32, 1 thread
M x N x K row: Accelerate develop PR MTGEMM-A col: Accelerate develop PR MTGEMM-A
512x512x512 1803 1600 1797 1813 1769 1272 1765 1778
1000x1000x1000 1844 1769 1812 1810 1764 1440 1756 1760
1024x1024x1024 1688 1553 1813 1812 1614 1412 1792 1800
2048x2048x2048 1687 677 1692 1698 1582 1330 1688 1694
3000x3000x3000 1570 526 1692 1687 1512 1345 1653 1654
4096x4096x4096 1574 1320 1680 1674 1525 1260 1675 1682
64x2112x7168 1177 477 1299 1346 908 462 1095 1107
64x24576x1536 637 395 982 989 1017 488 1101 1122
64x32768x512 627 353 945 962 872 421 989 1029
64x7168x16384 641 441 1030 1042 881 329 1053 1066
64x4096x7168 624 414 1056 1059 884 469 1076 1092
64x7168x2048 664 406 1026 1044 1001 377 1114 1157
128x2112x7168 1414 475 1559 1581 1223 720 1385 1390
128x24576x1536 947 390 1250 1273 1233 760 1399 1451
128x32768x512 920 358 1216 1235 1076 656 1288 1324
128x7168x16384 905 442 1303 1314 1157 561 1318 1318
128x4096x7168 915 413 1351 1351 1164 740 1357 1367
128x7168x2048 987 406 1309 1324 1221 606 1388 1420
4096x2112x7168 1594 1418 1797 1806 1452 1182 1669 1672
4096x24576x1536 1671 1359 1621 1660 1497 1225 1655 1670
4096x32768x512 1537 1503 1654 1664 1317 1133 1666 1663
4096x7168x16384 1507 1347 1680 1690 1515 1254 1644 1680
4096x4096x7168 1544 1312 1687 1681 1490 1278 1692 1685
4096x7168x2048 1701 1477 1667 1681 1535 1251 1661 1665
4096x256x4096 1438 1152 1562 1574 1190 832 1538 1555
11008x256x4096 1432 1158 1555 1562 1427 1099 1578 1598
4096x256x11008 1453 952 1603 1612 1141 843 1566 1558
5120x256x5120 1427 1171 1619 1628 1165 1091 1553 1552
13824x256x5120 1422 1176 1613 1620 1337 1118 1516 1528
5120x256x13824 1432 751 1578 1597 1170 1082 1542 1551
geometric mean, paper's 24 1127 707 1389 1404 1185 766 1391 1408
geometric mean, squares 512-4096 1691 1129 1747 1748 1624 1342 1721 1727
The paper's 24 workloads and squares 512-4096, fp32, default threads
M x N x K row: Accelerate develop PR MTGEMM-A col: Accelerate develop PR MTGEMM-A
512x512x512 2839 1662 3455 3481 2612 1273 3373 3417
1000x1000x1000 3603 1771 3676 3688 3399 1444 3547 3571
1024x1024x1024 3153 1571 3664 3718 2914 1425 3660 3652
2048x2048x2048 3405 667 3300 3308 3161 1333 3312 3325
3000x3000x3000 3290 528 3344 3339 3210 1345 3280 3281
4096x4096x4096 2841 1313 3252 3256 2947 1267 3262 3242
64x2112x7168 2172 472 2469 2524 1753 461 2203 2245
64x24576x1536 1355 388 1860 1870 2062 488 2103 2148
64x32768x512 1364 355 1738 1786 1727 424 1883 1999
64x7168x16384 1398 442 2030 2038 1877 329 2020 2039
64x4096x7168 1332 413 2064 2052 1914 473 2120 2158
64x7168x2048 1320 401 2043 2083 1939 371 2177 2275
128x2112x7168 2838 469 2967 3016 2588 724 2751 2783
128x24576x1536 2016 387 2410 2426 2640 756 2721 2833
128x32768x512 1976 357 2285 2333 2245 656 2486 2576
128x7168x16384 1880 441 2601 2622 2493 561 2590 2563
128x4096x7168 1811 412 2665 2638 2568 749 2691 2720
128x7168x2048 2112 402 2607 2619 2644 607 2716 2804
4096x2112x7168 3352 1415 3534 3552 2995 1180 3302 3317
4096x24576x1536 3683 1359 3277 3299 3248 1224 3260 3279
4096x32768x512 3452 1503 3272 3287 2790 1125 3260 3276
4096x7168x16384 3203 1360 3322 3345 3281 1252 3337 3323
4096x4096x7168 3181 1312 3281 3271 3136 1276 3288 3305
4096x7168x2048 3246 1479 3295 3334 2886 1253 3268 3274
4096x256x4096 3008 1152 3058 3081 2245 835 3081 3089
11008x256x4096 3045 1159 3037 3066 2769 1099 3018 3083
4096x256x11008 3135 908 3137 3177 2339 847 3101 3092
5120x256x5120 3048 1166 3133 3105 2281 1095 3128 3137
13824x256x5120 3039 1169 3133 3204 2793 1120 2940 2973
5120x256x13824 3086 734 3111 3154 2477 1073 3111 3102
geometric mean, paper's 24 2364 702 2709 2733 2445 766 2731 2770
geometric mean, squares 512-4096 3176 1135 3444 3460 3030 1346 3403 3411
The paper's 24 workloads and squares 512-4096, fp64, 1 thread
M x N x K row: Accelerate develop PR MTGEMM-A col: Accelerate develop PR MTGEMM-A
512x512x512 447 406 485 489 438 398 479 483
1000x1000x1000 487 406 474 461 447 404 469 457
1024x1024x1024 448 405 472 475 406 398 472 472
2048x2048x2048 399 362 453 453 389 349 452 452
3000x3000x3000 407 364 455 448 397 351 451 444
4096x4096x4096 394 346 432 446 388 339 439 447
64x2112x7168 318 228 417 415 326 224 345 351
64x24576x1536 217 147 365 368 304 232 358 365
64x32768x512 205 152 373 375 281 189 338 350
64x7168x16384 221 195 379 382 307 211 329 331
64x4096x7168 223 171 371 375 310 226 344 348
64x7168x2048 225 203 385 386 306 215 329 335
128x2112x7168 363 300 448 450 376 292 405 409
128x24576x1536 284 215 412 412 356 303 409 413
128x32768x512 265 229 421 421 326 259 400 406
128x7168x16384 290 256 416 422 356 284 391 394
128x4096x7168 295 242 416 417 356 297 399 403
128x7168x2048 295 266 425 426 359 290 393 395
4096x2112x7168 393 361 465 468 381 330 449 447
4096x24576x1536 389 359 451 455 374 339 443 454
4096x32768x512 384 410 462 465 343 339 447 452
4096x7168x16384 388 349 450 458 390 339 449 443
4096x4096x7168 389 343 436 448 397 340 448 443
4096x7168x2048 400 365 460 462 391 339 447 449
4096x256x4096 396 339 431 433 328 290 438 438
11008x256x4096 395 342 431 435 377 318 454 454
4096x256x11008 397 364 442 439 334 293 430 440
5120x256x5120 397 349 437 443 339 306 449 449
13824x256x5120 395 348 438 442 343 315 452 453
5120x256x13824 390 357 437 441 341 307 448 446
geometric mean, paper's 24 321 275 423 426 344 283 406 409
geometric mean, squares 512-4096 429 381 462 462 410 372 460 459
Small squares, fp32, 1 thread
M x N x K row: Accelerate develop PR MTGEMM-A col: Accelerate develop PR MTGEMM-A
4x4x4 4 1 1 1 5 <1 2 1
8x8x8 15 5 6 4 24 2 13 4
12x12x12 20 20 21 13 28 6 34 13
16x16x16 47 49 50 33 47 13 57 32
24x24x24 112 139 109 80 97 34 85 59
32x32x32 305 392 362 258 293 114 264 269
48x48x48 575 534 531 369 514 83 469 319
64x64x64 1018 941 980 1062 929 272 868 963
96x96x96 1370 1182 1316 1383 1255 438 1185 1253
128x128x128 1550 1361 1571 1633 1430 506 1424 1482
192x192x192 1683 1553 1724 1772 1597 708 1607 1651
256x256x256 1777 1643 1768 1804 1684 819 1695 1721
384x384x384 1814 1734 1870 1893 1787 1140 1805 1820
geometric mean, all 247 202 210 179 254 70 225 167
Small squares, fp32, default threads
M x N x K row: Accelerate develop PR MTGEMM-A col: Accelerate develop PR MTGEMM-A
4x4x4 4 <1 1 1 5 <1 2 1
8x8x8 15 5 6 4 24 2 13 4
12x12x12 21 20 21 14 28 6 33 13
16x16x16 48 49 49 33 47 13 56 32
24x24x24 107 138 107 79 97 34 85 63
32x32x32 313 396 351 274 311 111 244 270
48x48x48 573 534 533 366 515 83 473 320
64x64x64 1019 933 976 1073 929 272 861 956
96x96x96 1373 1148 1316 1381 1256 438 1165 1255
128x128x128 1549 1323 1568 1627 1433 505 1426 1484
192x192x192 1677 1543 1948 2043 1600 709 1712 1903
256x256x256 1764 1594 2927 2993 1693 817 2652 2867
384x384x384 1848 1729 3540 3560 1785 1138 3385 3425
geometric mean, all 248 190 230 199 256 70 243 185
Small squares, fp64, 1 thread
M x N x K row: Accelerate develop PR MTGEMM-A col: Accelerate develop PR MTGEMM-A
4x4x4 4 1 2 1 4 <1 3 1
8x8x8 11 2 12 5 16 2 14 5
12x12x12 17 8 21 12 21 7 27 10
16x16x16 42 42 41 27 43 40 40 24
24x24x24 89 28 87 62 87 22 63 52
32x32x32 192 77 164 139 187 76 149 135
48x48x48 323 128 216 188 294 126 175 175
64x64x64 388 174 360 389 354 168 330 358
96x96x96 437 199 400 403 406 199 377 379
128x128x128 457 232 442 457 428 227 422 436
192x192x192 471 275 462 474 455 266 444 458
256x256x256 474 337 461 479 456 333 451 466
384x384x384 486 379 482 488 476 375 475 479
geometric mean, all 123 56 112 89 124 51 111 83
Thin shapes, fp32, 1 thread
M x N x K row: Accelerate develop PR MTGEMM-A col: Accelerate develop PR MTGEMM-A
1x4096x4096 73 6 27 28 64 6 30 30
4x4096x4096 49 52 105 111 80 22 124 125
8x4096x4096 97 103 211 218 162 44 246 252
16x4096x4096 193 206 413 430 322 88 496 504
32x4096x4096 381 411 732 728 636 186 736 787
64x4096x4096 629 408 1044 1047 949 335 1022 1034
4096x1x4096 64 10 30 31 73 6 26 27
4096x4x4096 82 39 124 126 48 24 100 106
4096x8x4096 163 77 252 252 96 48 200 215
4096x16x4096 324 153 506 505 192 93 393 423
4096x32x4096 643 306 753 799 378 192 733 733
4096x64x4096 903 525 1029 1042 624 347 1048 1052
8x8x4096 48 64 95 95 48 6 95 95
16x16x4096 309 249 356 354 306 23 353 354
32x32x4096 741 897 1099 1127 739 131 1086 1127
geometric mean, all 198 118 268 273 197 49 263 270
Thin shapes, fp32, default threads
M x N x K row: Accelerate develop PR MTGEMM-A col: Accelerate develop PR MTGEMM-A
1x4096x4096 94 6 37 38 103 6 57 58
4x4096x4096 48 51 142 152 80 22 240 239
8x4096x4096 190 103 271 303 313 44 478 480
16x4096x4096 379 204 545 600 632 88 961 968
32x4096x4096 742 406 1402 1392 1240 186 1471 1535
64x4096x4096 1305 403 2065 2040 1861 335 2051 2074
4096x1x4096 103 10 58 58 94 6 41 43
4096x4x4096 82 39 241 240 48 24 154 158
4096x8x4096 314 77 482 484 189 48 289 302
4096x16x4096 624 153 969 973 375 92 625 609
4096x32x4096 1217 307 1489 1546 736 192 1408 1432
4096x64x4096 1893 526 2054 2092 1297 346 2050 2061
8x8x4096 48 64 95 96 49 5 95 95
16x16x4096 305 303 355 354 311 22 353 354
32x32x4096 731 819 1098 1125 739 133 1089 1128
geometric mean, all 298 118 412 422 298 49 421 427
Thin shapes, fp64, 1 thread
M x N x K row: Accelerate develop PR MTGEMM-A col: Accelerate develop PR MTGEMM-A
1x4096x4096 26 4 15 15 34 5 16 10
4x4096x4096 23 14 61 61 53 20 66 39
8x4096x4096 46 29 121 120 111 40 131 79
16x4096x4096 91 59 204 204 181 84 136 99
32x4096x4096 151 104 314 311 250 138 269 276
64x4096x4096 224 170 377 377 310 206 340 340
4096x1x4096 34 5 16 10 26 3 15 15
4096x4x4096 54 21 65 40 23 14 61 61
4096x8x4096 111 42 131 79 46 28 122 122
4096x16x4096 180 87 133 99 91 58 203 204
4096x32x4096 247 145 269 278 151 102 310 310
4096x64x4096 313 216 337 342 223 165 378 377
8x8x4096 84 12 95 95 84 12 95 95
16x16x4096 247 53 185 191 245 54 184 190
32x32x4096 334 92 380 387 331 92 378 387
geometric mean, all 103 40 125 112 103 39 126 112

SYMM, SYRK, SYR2K, TRMM, TRSM, per size

Column-major, lower, left side, no transpose, non-unit diagonal, beta 1; flops 2n^3 for GEMM, SYMM, SYR2K and n^3 for SYRK, TRMM, TRSM. Accelerate / develop / this PR, GFLOPS; the last column is this PR over Accelerate.

fp32, 1 thread
routine n Accelerate develop PR PR / develop PR / Accelerate
GEMM 64 908 265 751 2.8x 0.83x
GEMM 128 1411 497 1363 2.7x 0.97x
GEMM 256 1661 810 1652 2.0x 0.99x
GEMM 512 1725 1258 1738 1.4x 1.01x
GEMM 1024 1583 1377 1763 1.3x 1.11x
GEMM 2048 1554 1299 1659 1.3x 1.07x
GEMM 4096 1486 1235 1654 1.3x 1.11x
SYMM 64 687 32 212 6.6x 0.31x
SYMM 128 1169 80 473 5.9x 0.40x
SYMM 256 1495 104 796 7.7x 0.53x
SYMM 512 1620 116 993 8.6x 0.61x
SYMM 1024 1559 115 1318 11.5x 0.85x
SYMM 2048 1577 116 1356 11.7x 0.86x
SYMM 4096 1492 115 1437 12.5x 0.96x
SYRK 64 629 23 22 1.0x 0.03x
SYRK 128 1258 56 426 7.6x 0.34x
SYRK 256 1616 86 772 9.0x 0.48x
SYRK 512 1755 104 994 9.6x 0.57x
SYRK 1024 1758 108 1317 12.2x 0.75x
SYRK 2048 1496 112 1287 11.5x 0.86x
SYRK 4096 1480 113 1398 12.4x 0.94x
SYR2K 64 781 27 221 8.2x 0.28x
SYR2K 128 1336 61 509 8.3x 0.38x
SYR2K 256 1637 91 838 9.2x 0.51x
SYR2K 512 1746 106 1056 10.0x 0.60x
SYR2K 1024 1384 110 1329 12.1x 0.96x
SYR2K 2048 1473 113 1286 11.4x 0.87x
SYR2K 4096 1433 113 1405 12.4x 0.98x
TRMM 64 436 18 68 3.8x 0.16x
TRMM 128 905 55 204 3.7x 0.23x
TRMM 256 1299 85 452 5.3x 0.35x
TRMM 512 1492 108 729 6.8x 0.49x
TRMM 1024 1537 112 1107 9.9x 0.72x
TRMM 2048 1611 114 1238 10.9x 0.77x
TRMM 4096 1567 113 1273 11.3x 0.81x
TRSM 64 158 16 50 3.1x 0.32x
TRSM 128 362 41 105 2.6x 0.29x
TRSM 256 574 64 205 3.2x 0.36x
TRSM 512 736 85 353 4.2x 0.48x
TRSM 1024 791 99 584 5.9x 0.74x
TRSM 2048 867 106 822 7.8x 0.95x
TRSM 4096 997 109 1033 9.5x 1.04x
fp64, 1 thread
routine n Accelerate develop PR PR / develop PR / Accelerate
GEMM 64 349 163 305 1.9x 0.87x
GEMM 128 421 226 412 1.8x 0.98x
GEMM 256 444 327 449 1.4x 1.01x
GEMM 512 430 394 469 1.2x 1.09x
GEMM 1024 386 385 460 1.2x 1.19x
GEMM 2048 385 339 444 1.3x 1.15x
GEMM 4096 387 336 441 1.3x 1.14x
SYMM 64 286 30 141 4.7x 0.49x
SYMM 128 377 46 221 4.8x 0.59x
SYMM 256 420 54 286 5.3x 0.68x
SYMM 512 433 58 375 6.5x 0.87x
SYMM 1024 410 58 398 6.9x 0.97x
SYMM 2048 388 58 409 7.1x 1.05x
SYMM 4096 389 57 424 7.4x 1.09x
SYRK 64 283 19 18 0.9x 0.06x
SYRK 128 397 36 184 5.1x 0.46x
SYRK 256 439 48 255 5.3x 0.58x
SYRK 512 444 52 336 6.5x 0.76x
SYRK 1024 403 55 387 7.0x 0.96x
SYRK 2048 381 56 380 6.8x 1.00x
SYRK 4096 388 56 409 7.3x 1.05x
SYR2K 64 313 21 117 5.6x 0.37x
SYR2K 128 401 38 200 5.3x 0.50x
SYR2K 256 437 49 261 5.3x 0.60x
SYR2K 512 446 53 355 6.7x 0.80x
SYR2K 1024 391 55 374 6.8x 0.96x
SYR2K 2048 366 57 381 6.7x 1.04x
SYR2K 4096 368 56 411 7.3x 1.12x
TRMM 64 219 19 49 2.6x 0.22x
TRMM 128 337 37 113 3.1x 0.34x
TRMM 256 393 48 198 4.1x 0.50x
TRMM 512 408 55 295 5.4x 0.72x
TRMM 1024 433 57 358 6.3x 0.83x
TRMM 2048 410 57 377 6.6x 0.92x
TRMM 4096 407 57 402 7.1x 0.99x
TRSM 64 94 17 32 1.9x 0.34x
TRSM 128 155 30 61 2.0x 0.39x
TRSM 256 214 40 101 2.5x 0.47x
TRSM 512 231 49 176 3.6x 0.76x
TRSM 1024 265 53 257 4.8x 0.97x
TRSM 2048 286 55 312 5.7x 1.09x
TRSM 4096 320 55 364 6.6x 1.14x
fp32, default threads
routine n Accelerate develop PR PR / develop PR / Accelerate
GEMM 64 902 266 761 2.9x 0.84x
GEMM 128 1427 414 1365 3.3x 0.96x
GEMM 256 1675 815 2517 3.1x 1.50x
GEMM 512 2505 1200 3201 2.7x 1.28x
GEMM 1024 2885 1125 3600 3.2x 1.25x
GEMM 2048 2993 1246 3313 2.7x 1.11x
GEMM 4096 3146 1235 3257 2.6x 1.04x
SYMM 64 679 30 213 7.1x 0.31x
SYMM 128 1203 138 473 3.4x 0.39x
SYMM 256 1535 259 887 3.4x 0.58x
SYMM 512 2029 69 1278 18.5x 0.63x
SYMM 1024 2732 376 2350 6.2x 0.86x
SYMM 2048 2693 515 2773 5.4x 1.03x
SYMM 4096 2996 699 2806 4.0x 0.94x
SYRK 64 610 17 23 1.4x 0.04x
SYRK 128 1258 105 421 4.0x 0.33x
SYRK 256 1623 209 828 4.0x 0.51x
SYRK 512 1768 102 1284 12.6x 0.73x
SYRK 1024 2845 327 2003 6.1x 0.70x
SYRK 2048 1296 480 2088 4.3x 1.61x
SYRK 4096 2838 618 2450 4.0x 0.86x
SYR2K 64 755 66 223 3.4x 0.30x
SYR2K 128 1338 209 510 2.4x 0.38x
SYR2K 256 1641 229 924 4.0x 0.56x
SYR2K 512 1749 233 1331 5.7x 0.76x
SYR2K 1024 2811 224 1966 8.8x 0.70x
SYR2K 2048 2215 640 2033 3.2x 0.92x
SYR2K 4096 2835 797 2427 3.0x 0.86x
TRMM 64 436 6 68 11.3x 0.16x
TRMM 128 922 98 203 2.1x 0.22x
TRMM 256 1328 16 462 28.9x 0.35x
TRMM 512 1860 239 849 3.6x 0.46x
TRMM 1024 2643 203 1648 8.1x 0.62x
TRMM 2048 3052 460 2269 4.9x 0.74x
TRMM 4096 3196 752 2377 3.2x 0.74x
TRSM 64 158 1 50 50.0x 0.32x
TRSM 128 361 74 106 1.4x 0.29x
TRSM 256 573 18 203 11.3x 0.35x
TRSM 512 1070 194 401 2.1x 0.37x
TRSM 1024 1589 171 742 4.3x 0.47x
TRSM 2048 1866 501 1164 2.3x 0.62x
TRSM 4096 2159 722 1610 2.2x 0.75x
fp64, default threads
routine n Accelerate develop PR PR / develop PR / Accelerate
GEMM 64 350 97 304 3.1x 0.87x
GEMM 128 424 209 410 2.0x 0.97x
GEMM 256 449 323 757 2.3x 1.69x
GEMM 512 652 323 932 2.9x 1.43x
GEMM 1024 791 294 930 3.2x 1.18x
GEMM 2048 801 336 894 2.7x 1.12x
GEMM 4096 844 336 893 2.7x 1.06x
SYMM 64 286 3 141 47.0x 0.49x
SYMM 128 386 106 221 2.1x 0.57x
SYMM 256 433 15 354 23.6x 0.82x
SYMM 512 682 159 524 3.3x 0.77x
SYMM 1024 657 123 790 6.4x 1.20x
SYMM 2048 717 226 807 3.6x 1.13x
SYMM 4096 767 371 843 2.3x 1.10x
SYRK 64 284 12 19 1.6x 0.07x
SYRK 128 397 14 183 13.1x 0.46x
SYRK 256 440 111 291 2.6x 0.66x
SYRK 512 541 142 471 3.3x 0.87x
SYRK 1024 696 130 749 5.8x 1.08x
SYRK 2048 671 177 710 4.0x 1.06x
SYRK 4096 746 333 793 2.4x 1.06x
SYR2K 64 313 70 117 1.7x 0.37x
SYR2K 128 403 75 199 2.7x 0.49x
SYR2K 256 438 109 303 2.8x 0.69x
SYR2K 512 689 115 487 4.2x 0.71x
SYR2K 1024 681 172 666 3.9x 0.98x
SYR2K 2048 680 330 708 2.1x 1.04x
SYR2K 4096 702 410 794 1.9x 1.13x
TRMM 64 213 1 49 49.0x 0.23x
TRMM 128 344 5 112 22.4x 0.33x
TRMM 256 355 107 209 2.0x 0.59x
TRMM 512 704 106 365 3.4x 0.52x
TRMM 1024 820 132 626 4.7x 0.76x
TRMM 2048 833 310 715 2.3x 0.86x
TRMM 4096 864 403 775 1.9x 0.90x
TRSM 64 94 2 32 16.0x 0.34x
TRSM 128 155 6 61 10.2x 0.39x
TRSM 256 209 12 109 9.1x 0.52x
TRSM 512 408 115 215 1.9x 0.53x
TRSM 1024 561 127 373 2.9x 0.66x
TRSM 2048 442 322 488 1.5x 1.10x
TRSM 4096 706 402 636 1.6x 0.90x

GEMM per call, small squares

Nanoseconds per call, one thread, minimum over many short batches in three passes (bench_small); row-major beta 0, column-major beta 1. A ratio above 1 means this PR is faster.

row-major, fp32
n develop PR MTGEMM-A develop / PR MTGEMM-A / PR
4 141 142 229 0.99 1.61
8 158 137 248 1.15 1.81
12 137 137 261 1.00 1.91
14 139 140 255 0.99 1.82
15 139 146 266 0.95 1.82
16 136 137 247 0.99 1.80
17 213 214 331 1.00 1.55
18 206 215 323 0.96 1.50
20 190 192 340 0.99 1.77
24 171 171 341 1.00 1.99
28 186 180 386 1.03 2.14
32 135 138 237 0.98 1.72
36 323 323 448 1.00 1.39
40 338 338 484 1.00 1.43
44 359 359 531 1.00 1.48
48 375 375 573 1.00 1.53
52 469 469 859 1.00 1.83
56 479 479 943 1.00 1.97
60 516 521 1036 0.99 1.99
64 516 500 490 1.03 0.98
68 1031 859 995 1.20 1.16
72 1057 885 1036 1.19 1.17
76 1109 932 1094 1.19 1.17
80 1130 922 1109 1.23 1.20
84 1312 1266 1677 1.04 1.32
88 1422 1401 1844 1.01 1.32
92 1495 1380 1870 1.08 1.36
96 1526 1318 1271 1.16 0.96
100 2500 1922 2135 1.30 1.11
104 2536 1932 2177 1.31 1.13
108 2688 2078 2318 1.29 1.12
112 2688 2005 2344 1.34 1.17
116 3000 2661 3276 1.13 1.23
120 3021 2661 3380 1.14 1.27
124 3062 2781 3505 1.10 1.26
128 3114 2604 2568 1.20 0.99
132 4625 3625 3708 1.28 1.02
136 4708 3666 3750 1.28 1.02
140 5375 3833 3916 1.40 1.02
144 5375 3833 3875 1.40 1.01
148 5833 4833 5375 1.21 1.11
152 5875 4833 5416 1.22 1.12
156 5583 5000 5583 1.12 1.12
160 6000 4958 4709 1.21 0.95
164 7958 6375 6458 1.25 1.01
168 8542 6416 6542 1.33 1.02
172 8333 6666 6833 1.25 1.03
176 8750 6541 6791 1.34 1.04
180 9416 8000 8667 1.18 1.08
184 9458 8083 8833 1.17 1.09
188 9666 8375 9166 1.15 1.09
192 9459 8041 7875 1.18 0.98
224 14292 12666 12333 1.13 0.97
256 20167 18708 18458 1.08 0.99
row-major, fp64
n develop PR MTGEMM-A develop / PR MTGEMM-A / PR
4 155 38 214 4.08 5.63
8 193 73 217 2.64 2.97
12 412 142 275 2.90 1.94
14 425 214 295 1.99 1.38
15 451 266 296 1.70 1.11
16 183 186 297 0.98 1.60
17 763 281 363 2.72 1.29
18 789 284 373 2.78 1.31
20 933 280 391 3.33 1.40
24 956 301 435 3.18 1.45
28 1136 380 577 2.99 1.52
32 764 400 467 1.91 1.17
36 1635 557 714 2.94 1.28
40 1771 588 766 3.01 1.30
44 2000 958 1062 2.09 1.11
48 1531 1005 1146 1.52 1.14
52 3141 1203 1578 2.61 1.31
56 3354 1271 1688 2.64 1.33
60 3589 1495 6630 2.40 4.43
64 2844 1438 1344 1.98 0.93
68 5162 2047 3411 2.52 1.67
72 5672 2099 3536 2.70 1.68
76 6193 3031 5854 2.04 1.93
80 4792 3146 6042 1.52 1.92
84 8625 3614 8708 2.39 2.41
88 9068 3734 8932 2.43 2.39
92 9693 4234 12203 2.29 2.88
96 8229 4370 4385 1.88 1.00
100 12281 5302 7573 2.32 1.43
104 12760 5432 7812 2.35 1.44
108 13719 7318 11312 1.87 1.55
112 12010 7490 11552 1.60 1.54
116 17135 8286 15573 2.07 1.88
120 17927 8526 15911 2.10 1.87
124 18750 9359 22156 2.00 2.37
128 17458 9364 9125 1.86 0.97
132 20041 11166 14125 1.79 1.27
136 21000 11416 14417 1.84 1.26
140 21667 14416 20000 1.50 1.39
144 21125 14750 20333 1.43 1.38
148 26166 16000 26208 1.64 1.64
152 28750 16333 26666 1.76 1.63
156 28292 17666 33166 1.60 1.88
160 27916 18041 17875 1.55 0.99
164 35542 20416 24542 1.74 1.20
168 35292 20750 24958 1.70 1.20
172 38416 25334 32125 1.52 1.27
176 36375 25750 32583 1.41 1.27
180 44250 27625 40125 1.60 1.45
184 45458 28125 40625 1.62 1.44
188 47000 30041 52125 1.56 1.74
192 45875 30208 29625 1.52 0.98
224 68209 47708 47458 1.43 0.99
256 96333 70834 69625 1.36 0.98
col-major, fp32
n develop PR MTGEMM-A develop / PR MTGEMM-A / PR
4 453 31 241 14.61 7.77
8 467 48 255 9.73 5.31
12 535 75 259 7.13 3.45
14 618 114 263 5.42 2.31
15 665 152 271 4.38 1.78
16 603 114 251 5.29 2.20
17 688 180 414 3.82 2.30
18 699 195 416 3.58 2.13
20 712 385 425 1.85 1.10
24 772 273 436 2.83 1.60
28 850 292 478 2.91 1.64
32 546 241 240 2.27 1.00
36 1974 396 495 4.98 1.25
40 2078 406 552 5.12 1.36
44 2635 464 625 5.68 1.35
48 2375 453 656 5.24 1.45
52 2682 672 1031 3.99 1.53
56 2740 693 1125 3.95 1.62
60 3000 766 1208 3.92 1.58
64 1703 578 547 2.95 0.95
68 4588 1021 1088 4.49 1.07
72 4818 1036 1120 4.65 1.08
76 5927 1088 1193 5.45 1.10
80 5276 1078 1198 4.89 1.11
84 5974 1562 1943 3.82 1.24
88 6068 1688 2104 3.59 1.25
92 6594 1688 2115 3.91 1.25
96 3833 1474 1396 2.60 0.95
100 8672 2193 2318 3.95 1.06
104 9172 2224 2401 4.12 1.08
108 10120 2380 2562 4.25 1.08
112 9620 2307 2583 4.17 1.12
116 10724 3135 3708 3.42 1.18
120 10557 3135 3812 3.37 1.22
124 11547 3302 3964 3.50 1.20
128 7391 2885 2823 2.56 0.98
132 12834 4291 4333 2.99 1.01
136 13208 4291 4333 3.08 1.01
140 14292 4500 4584 3.18 1.02
144 13875 4208 4208 3.30 1.00
148 15333 5750 6250 2.67 1.09
152 15375 5708 6291 2.69 1.10
156 16583 6125 6708 2.71 1.10
160 11125 5333 5125 2.09 0.96
164 18875 7208 7250 2.62 1.01
168 20083 7250 7375 2.77 1.02
172 21959 7625 7750 2.88 1.02
176 20375 7208 7375 2.83 1.02
180 21958 9208 9916 2.38 1.08
184 21334 9291 10000 2.30 1.08
188 23750 9708 10416 2.45 1.07
192 17208 8666 8416 1.99 0.97
224 23792 13500 13125 1.76 0.97
256 34666 19625 19333 1.77 0.99
col-major, fp64
n develop PR MTGEMM-A develop / PR MTGEMM-A / PR
4 219 32 225 6.84 7.03
8 426 56 222 7.61 3.96
12 462 104 312 4.44 3.00
14 502 159 332 3.16 2.09
15 518 202 335 2.56 1.66
16 188 190 316 0.99 1.66
17 1329 374 452 3.55 1.21
18 1155 377 507 3.06 1.34
20 1242 395 520 3.14 1.32
24 1239 434 536 2.85 1.24
28 1439 525 818 2.74 1.56
32 789 430 487 1.83 1.13
36 2396 646 745 3.71 1.15
40 2359 672 802 3.51 1.19
44 2734 1208 1172 2.26 0.97
48 1458 1234 1229 1.18 1.00
52 4266 1510 1766 2.83 1.17
56 4010 1568 1828 2.56 1.17
60 4719 1864 7021 2.53 3.77
64 2995 1568 1464 1.91 0.93
68 6943 2260 3583 3.07 1.59
72 6656 2323 3729 2.87 1.61
76 7760 3531 6135 2.20 1.74
80 5083 3651 6359 1.39 1.74
84 11146 4151 9177 2.69 2.21
88 10411 4297 9474 2.42 2.20
92 11714 4844 12656 2.42 2.61
96 8453 4635 4620 1.82 1.00
100 15250 5745 7912 2.65 1.38
104 14510 5844 8135 2.48 1.39
108 16521 8151 11870 2.03 1.46
112 12417 8266 12057 1.50 1.46
116 20682 9167 16182 2.26 1.77
120 19620 9375 16604 2.09 1.77
124 21901 10307 22859 2.12 2.22
128 17917 9844 9604 1.82 0.98
132 24416 12041 14875 2.03 1.24
136 23125 12083 14875 1.91 1.23
140 26750 15833 20750 1.69 1.31
144 20667 15916 20875 1.30 1.31
148 31166 17416 27125 1.79 1.56
152 29000 17583 27416 1.65 1.56
156 32666 19208 34333 1.70 1.79
160 27708 18750 18583 1.48 0.99
164 40667 21666 25500 1.88 1.18
168 38959 21666 25833 1.80 1.19
172 42292 27291 33166 1.55 1.22
176 37209 27291 33333 1.36 1.22
180 51042 29791 41708 1.71 1.40
184 50417 29750 42333 1.69 1.42
188 55708 32083 53208 1.74 1.66
192 46792 31333 30708 1.49 0.98
224 69458 49084 48875 1.42 1.00
256 98416 72833 71625 1.35 0.98

Credit

Design: Deng et al. (arXiv:2512.21473). MTGEMM-A is the measured reproduction this port comes from. The
existing kernels of #5971 (ported from vlovero's ARMv9.2-GEMM) and the SME1 direct kernels (#5084, #5380) remain
in place as fallbacks and for small problems.

Prepared with AI coding assistance.

On SME2 cores with a 512-bit streaming vector length (checked at run time), sme_sgemm_kernel and
sme_dgemm_kernel run a blocked SME2 GEMM after Deng et al., "Demystifying ARM SME to Optimize General Matrix
Multiplications" (arXiv:2512.21473), as measured in MTGEMM-A: blocking from the paper's analytical model, the ZA
transposition of A in packing, online packing of B, a 1 x NT-tile micro-kernel with multi-vector loads plus
half-width and edge kernels, four-row ZA moves for C, core-side L2 prefetch, all four transpose cases, and one
thread per SME unit through exec_blas from 2^22 multiply-adds. The existing kernels stay as the fallback.

- #pragma clang attribute enables SME2 for these functions only (Apple clang >= 17, LLVM >= 18); no build
  flags change.
- Small problems: a per-thread packing buffer; interface/gemm.c keeps problems below 20^3 (fp64 16^3)
  multiply-adds on NEON; the SME1 direct sgemm keeps row-major problems below 64^3.
- M, N or K of 2^30 or more stay on the existing kernels; a failed blas_memory_alloc falls back to one thread
  and malloc.
- SME units: one per performance cluster on Apple; two per Super cluster on M5 Pro/Max.
- cpuid: detect the M4 Pro (hw.cpufamily 0x17d5b93a) and later Apple cores with FEAT_SME as VORTEXM4.
For real single and double precision on SME targets, kernel/arm64/sme_level3.h splits the symmetric or
triangular dimension recursively: off-diagonal blocks are GEMM calls on SME_[SD]GEMM_KERNEL, diagonal blocks of
at most 64 run as one GEMM on a full copy (SYMM, SYRK, SYR2K, TRMM), and TRSM solves blocks of at most 32 by
substitution in a local tile (NEON transposes for the left side). The interface files only call the hooks, after
the argument checks and develop's SME1 direct kernels, from 2e5 multiply-adds; the level-3 driver and its
kernels are unchanged. TRMM falls back to the driver when no buffer is available.
@tesch1

tesch1 commented Sep 30, 2026

Copy link
Copy Markdown
Author

Superseded by #6074, which does the same work without changes to the interface files.

This PR ran SYMM, SYRK, SYR2K, TRMM and TRSM on the SME GEMM kernel through hooks in interface/{symm,syrk,syr2k,trsm}.c, which included a header from kernel/arm64/, and it added an M4-tuned size threshold to interface/gemm.c. #6074 adds SME2 kernels for the level-3 driver instead (KERNEL.VORTEXM4 and param.h), so the LAPACK routines on the driver also run on SME2. It keeps the small-size limit inside the kernel, and it limits level-3 threads to one per SME unit in num_cpu_avail(). #6074 also compares the speed of both PRs; this PR stays faster for some fp64 routines and for very small SGEMMs.

@tesch1 tesch1 closed this Sep 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

arm64 SME2 (Apple M4): SGEMM/DGEMM well below the hardware, level-3 routines on NEON, M4 Pro not detected

1 participant