Methodology — the validation toolkit¶
A SIMD kernel is not "done" because it compiles and is fast on one machine. In go-simd it is done when it is proven correct against the reference it replaces and measured honestly on representative hardware. Every repo in the organization follows the same five-stage pipeline.
1. Check existing work first¶
Before a single instruction is written, the field is surveyed so the comparison is against the real state of the art, not a strawman. Concretely, the prior pure-Go SIMD implementations each repo benchmarks against:
| repo | prior art benchmarked / surveyed |
|---|---|
| base64 | emmansun/base64 (AVX2), cristalhq/base64 (scalar); aklomp/base64 is cgo (excluded) |
| base32 | none — no pre-existing pure-Go SIMD base32 exists |
| hex | tmthrgd/go-hex (SSE/AVX, archived Sep 2025) |
| utf8 | stuartcarnie/go-simd (2018 Lemire port); charlievieth/simdutf is cgo (excluded) |
| adler32 | mhr3/adler32-simd (Chromium/zlib transpiled via gocc) |
| popcount | barakmich/go-popcount (Tom Thorogood asm) |
| histogram | vteromero/byte-hist (CLI, not a library) — gap filled |
| bitpack | ronanh/intcomp (head-to-head is a follow-up) |
| matchlen | the compressor in-tree primitives (LZ4/zstd match-finders) |
| strconv | Go's own strconv (already very tight, especially Atoi's inlined path) |
This is why the headlines are credible: each is a measured result against a named competitor, and cgo wrappers (faster, but needing a C toolchain) are explicitly excluded from the pure-Go tier rather than quietly ignored.
2. Generate the assembly with go-asmgen (6 architectures)¶
No kernel is hand-encoded. The .s is emitted by
go-asmgen v0.5.0 — one builder over a
shared ABI0 layout for all six of Go's 64-bit SIMD targets — and
cmd/asm does the encoding:
| arch | ISA used by go-simd kernels |
|---|---|
| amd64 | SSE2 / SSSE3 / SSE4.1 + AVX2, runtime-dispatched via x/sys/cpu |
| arm64 | NEON |
| loong64 | LSX / LASX |
| riscv64 | RVV |
| ppc64le | VSX / AltiVec (POWER8 baseline, no runtime dispatch) |
| s390x | vector facility (z13 baseline) — big-endian |
This is the go-asmgen v0.5.0 six-arch milestone: one generator now emits correct kernels for every 64-bit target Go can produce SIMD for. The ppc64le (VSX) and s390x (vector facility) backends are the newest; since VSX is baseline on POWER8+ and the vector facility on z13+, those kernels need no runtime feature dispatch — the SIMD path is simply the arch's only path (build-tagged), like riscv64 and loong64.
s390x is big-endian. Every kernel is ported and validated on a big-endian
target, which exercises the control-vector lane numbering of VPERM shuffles and
the byte order of vector loads/stores (VL/VST) — a real cross-endian
correctness check, not just a recompile. Where amd64's 16-bit windows are
themselves big-endian (e.g. base32, base64), the s390x layout matches directly.
The *_gen.go generators are //go:build ignore (go-asmgen is a build-time
tool, not a runtime dependency); the resulting .s is committed. Constant
tables are emitted via go-asmgen's emit.File.Data. Architectures without a
kernel use a portable scalar fallback so the package builds and is correct
everywhere.
3. Model the cycles with llvm-mca¶
Before benchmarking, the inner loop is given a port-level cycle model with
llvm-mca (go-asmgen toolkit/mca).
This is what turned guesses into wins:
- base64 —
llvm-mcashowed the kernel was shuffle-port (Zn3FP1) bound; disassemblingemmansunrevealed a-4-offset 32-byte load that spreads with a singleVPSHUFB(no cross-laneVINSERTI128). Combining it with the 2× unroll was predicted at 2.15 cyc/block (vs emmansun's 2.20) and the benchmark confirmed it — the ~5–6% lead. - adler32 — the 4×-unrolled AVX2 kernel models at 17.9 bytes/cycle
(≈99% of the vector-ALU port ceiling) on the Zen3 model. The static model puts
it at parity with
mhr3; the real out-of-order frontend still leaves mhr3 ~7% ahead. The model is a guide, not a verdict — see stage 4.
4. Validate on real hardware (Rosetta is never trusted for AVX2)¶
Correctness is proven by differential tests + fuzzing against the stdlib/oracle reference, and the headline benchmarks come from native CI, not emulation:
- Real AVX2 x86-64 host for amd64. Rosetta hides AVX2, so it is never
used to validate or benchmark amd64 kernels; a
--platform amd64container crashes Go. A real x86-64 box (or KVM guest) and GitHub Actionsubuntu-latest(AMD EPYC, AVX2) are the authoritative amd64 sources. - Native arm64 (the Apple-silicon dev box) for the NEON kernels and the
!amd64generic fallback. - Real silicon on the GCC Compile Farm for the non-x86/arm SIMD targets. As of 2026-06, ppc64le (POWER9, plus POWER8E for the ISA-baseline fallback), riscv64 (SpacemiT X60, RVV 1.0) and loong64 (Loongson 3A5000, LSX) are run and benchmarked natively — superseding the earlier qemu-only / llvm-mca estimates on those arches. The portable scalar fallback is additionally build+test-validated on a seventh architecture, ppc64 (big-endian), on real POWER9.
- QEMU is still the correctness lane for the targets without native silicon,
and the cross-check everywhere: riscv64 (RVV), loong64 (LSX), ppc64le (VSX)
and s390x (vector facility) run under
qemu-userin a debian:trixie container (QEMU_CPU=power9/qemu). QEMU's TCG does not model out-of-order execution, so emulated MB/s is not representative and is never quoted as a headline. - s390x is now natively measured on real IBM z15 (VXE2, native execution,
2026-07-03,
-count=6). It runs the official vectors and the byte-identical differential fuzz (and on big-endian s390x the output is proven bit-exact), and the vector-facility kernels are now measured on real IBM Z silicon — no longer extrapolated from emulation. Six SIMD targets, validated on seven architectures. - Fuzzing runs against the stdlib reference on arbitrary input, comparing
the returned value and the full error (e.g.
strconv's*NumError,hex'sInvalidByteErroroffset). Direct kernel fuzz targets exercise the SSE and AVX2 paths separately (millions of executions, zero mismatches).
5. 100% statement coverage, enforced as a CI gate¶
Every repo gates 100% Go statement coverage in CI, on every arch job — the build fails below 100%. Force tests drive every dispatch branch (AVX2 / SSE / POPCNT / scalar fallback) directly on the native runner, regardless of what the runtime CPU would otherwise select, so no branch is left unmeasured.
What coverage measures
The figure is of the Go code only. The generated .s SIMD kernels are
not measured by go test -cover; they are validated by the differential
tests against the stdlib/oracle reference plus the fuzz targets. "100%
coverage" therefore means every Go statement, including every fallback
branch — and the assembly is covered by the correctness suite, not the
coverage counter.
Why this matters¶
The combination — one generator across six ISAs, a cycle model that predicts
before it measures, real-hardware validation that refuses to trust emulation (and
qemu-correctness validation where no native runner exists), and a coverage gate
that refuses to ship an unexercised branch — is what lets go-simd publish
honest numbers: it beats emmansun and tmthrgd where it says it does, it
openly reports the ~7% gap to mhr3 and the fact that a byte histogram has no
SIMD win at all, and it quotes ppc64le/s390x throughput only from real hardware —
s390x on real IBM z15, ppc64le on real POWER9 where a runner is available —
rather than quoting emulated throughput as if it were a headline. A standout of the six-arch port: base32 gets real SIMD on ppc64le
(VSRH, register-variable vector shift) and s390x (VMLHH, integer vector
multiply-high) precisely where Go's arm64 assembler lacks those two
primitives — so POWER and IBM Z run the full spread-extract kernel that NEON
could not express.