strawmANN benchmark report

strawmANN against Qdrant · AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical) · generated 2026-09-29 20:15:09 CEST

Dataset
sift1m
Vectors
1,000,000 × 128 · 10,000 held-out queries
Metric
euclid
Storage datatype
strawmANN float32 · Qdrant float32 (default)
Quantization
none (fp32) · separate rows measure binary (1 bit/dim), PQ (product), SQ8 scalar
strawmANN: 43/43 ok Qdrant: 43/43 ok

Summary

At equal recall, strawmANN serves 1.81x to 2.16x Qdrant's throughput. That range spans the recall levels measured; allowing for the uncertainty in each recall figure widens it to 0.86x to 5.14x.

throughput at equal recall
1.81 to 2.16x
strawmANN over Qdrant, recall@10 0.909 to 0.999
p99 latency at 90% load
strawmANN1.41 ms
Qdrant5.27 ms
W4-sat90, open loop, one offered rate
upload and index build
strawmANN42 s
Qdrant80 s
W1 + W2
peak memory (RSS)
strawmANN5.2 GiB
Qdrant5.2 GiB
largest before the concurrent-write rows (W11); of it, anonymous 2.1 GiB / 833.8 MiB
written to disk
strawmANN2.5 GiB
Qdrant12.2 GiB
before the concurrent-write rows (W11); they wrote 122.6 MiB / 10.1 GiB more

Where strawmANN is slower, past the noise floor:

Throughput against recall

Read it vertically: at any recall both engines reach, the higher curve is faster. Throughput from bfb W10, recall from the conformance sweep, joined on ef. ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.

Every figure is the median of 3 passes per engine, run alternately (A/B/A/B). What this does not establish.

Glossary
recall@10
Of the ten nearest neighbours a query really has, the share the engine returned. 1.0 is a perfect answer. "Really has" is settled by an exhaustive fp64 search, not by the other engine.
ef
The candidate-list size: how many nodes the index keeps in play while searching. Larger is slower and more accurate. It does not mean the same amount of work in both engines, which is why the headline holds recall fixed instead.
MRDE
Mean relative distance error: when the engine returns a wrong neighbour, how much further away it is than the right one. Small numbers mean the misses were near-misses.
qps
Queries per second. On a batched row one request carries several queries, and the request rate is shown under it.
p50 / p99 / p99.9
The latency half of requests beat, that 99% beat, that 99.9% beat. The tail is what a user notices.
closed / open loop
A closed loop sends the next request only when the last one comes back, so a slow server receives less work and its tail looks better than it is. An open loop sends at a fixed rate regardless.
noise floor
How much a number moves between identical runs on this machine. A ratio inside it is shown grey with ≈: no measured difference. Hover a ratio to see its band.
conformance tier
What the two engines were shown to agree on before any speed was quoted. T1 licenses a single engine's own numbers; T3 licenses comparing the two, and requires their recall to be statistically indistinguishable.
§ numbers
Sections of docs/spec.md, the written rule each claim in the appendix is measured against. §7.1 is the host gate, §7.4 the comparison rules, §8 what may be published.

Throughput

Queries per second; higher is better. The ratio is strawmANN over Qdrant: green is faster, red slower, grey ≈ inside the noise floor (hover a ratio for its band). A dash means the pair is not compared, and the note says why. An ef sweep is one row showing its range; every measurement is under All throughput rows.

workloadstrawmANNQdrantrationotes
W3search, fp32, single query1,7022,0080.85x
W4search, saturating (closed loop)23,05914,1351.63x
W5search batched (16 distinct dataset queries per request)23,4001,437 requests/s4,859304 requests/s4.82x
W6quantized: scalar4,3934,6900.94x
W6 ef 32 to 512SQ8 recall control (latency only)1,835 to 9,3751,830 to 7,7540.94x3 of 5 points at parity 1 of 5 points not compared recall differs
W7quantized: binary + oversampling4,5084,655-recall below the usable floor: 0.0618 and 0.0660, both under 0.50 — this row characterises the encoding, not the engines
W8quantized: PQ4,4904,1651.08x
W9exact / brute force1861241.50x
W10 ef 32 to 512recall control (latency only)6,622 to 40,8203,963 to 22,0221.67 to 2.03x1 of 5 points not compared recall differs
W12-sel1filtered search, one keyword (~1% of bench12)5,8073,8841.49x
W12-sel10filtered search, any of 10 keywords (~10% of bench12)1,3592,222-recall differs: 1.0000 vs 0.9893
W12-sel1 ef 32 to 512filtered recall control, one keyword (latency only)5,996 to 6,4191,701 to 6,4321.19 to 3.66x1 of 5 points at parity
W12-sel10 ef 32 to 512filtered recall control, any of 10 keywords (latency only)1,358 to 3,172798 to 4,6851.02 to 1.70x3 of 5 points not compared recall differs
W13scroll / pagination51,19246,5601.10xstrawmANN: 92% is client and socket, not server Qdrant: 82% is client and socket, not server
W11-steadymixed read/write below the rebuild threshold: search bench2 while 50,000 synthetic points append18,2285,537-search-during-write; no recall join
W11mixed read/write: search bench2 while 200,000 synthetic points append (runs last)8,0333,800 requests/s2,387-search-during-write; no recall join

Throughput by workload

Higher is better. Hatched bars are rows the table does not compare (W7, W11): each bar is that engine's own rate, and the pair is not a result.

Recall

recall@10 against an exact fp64 search, not against the other engine. Equal ef is not equal work in the two engines, so the comparison that counts holds recall fixed and compares throughput there.

At matched recall

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.9089strawmANN at ef=32, Qdrant interpolated40,82020,7681.97x1.88 – 2.06x
0.9603Qdrant at ef=64, strawmANN interpolated31,95815,4942.06x2.02 – 2.11x
0.9639strawmANN at ef=64, Qdrant interpolated31,42514,6652.14x1.93 – 2.38x
0.9875Qdrant at ef=128, strawmANN interpolated21,06610,1592.07x1.99 – 2.16x
0.9888strawmANN at ef=128, Qdrant interpolated20,6089,5532.16x1.68 – 2.77x
0.9967Qdrant at ef=256, strawmANN interpolated12,6226,5911.91x1.76 – 2.08x
0.9974strawmANN at ef=256, Qdrant interpolated12,0895,7472.10x0.86 – 5.14x
0.9992Qdrant at ef=512, strawmANN interpolated7,1623,9631.81x1.51 – 2.17x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W10 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column is why the ratio is not a result on its own. The ratio treats the anchor's recall as exact; it is an estimate, and moving it across its 95% interval moves the interpolated rate with it — near the top of the sweep 0.003 of recall spans a factor of 1.8, so two decimals there quote the interpolation. Even that is the narrow reading: it moves one recall and not the other, adjacent anchors share bracketing segments, and the throughputs behind it are medians of the run's passes.

At matched recall, SQ8

The same reading over the scalar-quantized collection. Under the `pool` policy both engines rescore `ef` candidates, so W6 is compared row by row as well; held at equal recall, the reading does not depend on the rows landing in one recall band. Note where each engine's curve stops, because a recall only one of them reaches is the more useful fact about an encoding than any ratio.

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.9087strawmANN at ef=32, Qdrant interpolated9,3757,4611.26x1.23 – 1.28x
0.9600Qdrant at ef=64, strawmANN interpolated6,5786,1791.06x1.03 – 1.10x
0.9646strawmANN at ef=64, Qdrant interpolated6,3725,8961.08x1.04 – 1.12x
0.9871Qdrant at ef=128, strawmANN interpolated4,5354,6840.97x0.93 – 1.00x
0.9890strawmANN at ef=128, Qdrant interpolated4,4114,3211.02x0.93 – 1.12x
0.9968Qdrant at ef=256, strawmANN interpolated3,0903,0691.01x0.95 – 1.07x
0.9974strawmANN at ef=256, Qdrant interpolated3,0132,7241.11x0.88 – 1.39x
0.9993Qdrant at ef=512, strawmANN interpolated1,9691,8301.08x0.93 – 1.24x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W6 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column reads as it does in the fp32 table above.

At matched recall, filtered to 10%

The same reading under a keyword filter over the points the condition matched, 19,882 to 20,160 across 3 builds in strawmANN and 19,933 to 20,042 across 3 builds in Qdrant (bfb draws the keyword payloads unseeded at each upload). The per-row table refuses W12-sel10 a ratio because the two engines land just outside the recall band at the one ef it measures; held at equal recall instead, the comparison exists at every recall both engines reach. Note the shape rather than any single number: one engine's curve is flat in ef and the other's is steep, so where you match decides the ratio.

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.9777strawmANN at ef=32, Qdrant interpolated3,1722,4271.31x1.27 – 1.34x
0.9893Qdrant at ef=128, strawmANN interpolated2,4172,2261.09x1.03 – 1.15x
0.9954strawmANN at ef=64, Qdrant interpolated2,0981,6161.30x1.21 – 1.40x
0.9992Qdrant at ef=256, strawmANN interpolated1,4701,3251.11x1.04 – 1.18x
1.0000Qdrant at ef=512, strawmANN interpolated1,3647981.71x1.66 – 1.76x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W12-sel10 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column reads as it does in the fp32 table above.

Recall by ef

efstrawmANNQdrant
recall@1recall@10recall@100MRDErecall@1recall@10recall@100MRDE
ef=320.94550.9089-3.14e-030.93600.8986-3.58e-03
ef=640.98020.9639-1.13e-030.97710.9603-1.25e-03
ef=1280.99540.98880.95093.28e-040.99360.98750.94403.34e-04
ef=2560.99940.99740.98551.32e-040.99830.99670.98357.66e-05
ef=5120.99990.99950.99689.10e-060.99950.99920.99612.26e-05
How recall was measured

10,000 held-out queries, limit 10, ε=4.172e-07, against the fp64 oracle rather than against the other engine. Base checksum 7bf7ead7809a45cd: the same corpus the latency rows were measured on. Each pass rebuilds the collection. Across 3 independent builds of it, recall@10 at ef=512 spread by: strawmANN 0.00611; Qdrant 0.00017. strawmANN's graphs disagree with each other 36x as much as Qdrant's do, so a ratio read at the top of this table reports which graph the build produced as much as which engine built it. Over the same builds the engine reported unreachable nodes: strawmANN 0 of 200,000 and 0 of 1,000,000 — a node nothing points at is invisible at any ef, so that is where a graph-quality difference shows rather than being inferred from the recall beside it. Level seed: strawmANN 0x57ea3111. Recorded as provenance: under this engine's level draw, six builds at four seeds sit within 0.00003 of recall@10 at ef 512, well inside the spread above.

What each quantization costs in recall

Colour is the encoding here, not the engine — the engines are the line style. Each sweep is joined on the quantization parameters it actually sent, so a curve speaks only for the search its row ran. Higher is better, and the encodings are not free: read this against the throughput their rows bought.

Latency

Client-side round trip. A closed loop understates the tail (a stalled server stops receiving requests), so this keeps the fixed-rate, open-loop rows, which offer the same load to both engines, plus single-client W3 and batched W5. Every row is under All latency rows.

workloadstrawmANNQdrant
p50p99p99.9p50p99p99.9
W3 search, fp32, single query598 µs715 µs792 µs505 µs594 µs666 µs
W4-sat50 search, fixed rate at 50% of saturation (open loop)493 µs1.01 ms1.56 ms753 µs1.83 ms2.32 ms
W4-sat70 search, fixed rate at 70% of saturation (open loop)532 µs1.12 ms1.22 ms838 µs2.02 ms2.87 ms
W4-sat90 search, fixed rate at 90% of saturation (open loop)718 µs1.41 ms1.53 ms1.24 ms5.27 ms9.40 ms
W5 search batched (16 distinct dataset queries per request)1.36 ms1.52 ms2.63 ms6.56 ms7.10 ms7.74 ms

Ingest and index build

Qdrant indexes while it ingests, so its upload time already contains most of the indexing; strawmANN uploads raw and builds afterwards. Compare the sum, not the upload line.

Where the load time goes

Solid is upload, hatched is the wait for Green. The bar's whole length is the sum this section asks you to compare; the split is why the upload line alone is not comparable between these engines.

workloadstrawmANNQdrant
W0-uploadd=4 floor: graph traversal with the distance taken out19.06 s26.06 s
W1ingest throughput (no index wait)1.08 s11.83 s
W2index build, time to Green41.13 s68.16 s
W6-uploadsearch, SQ8 scalar quantization41.12 s53.12 s
W7-uploadsearch, binary quantization41.13 s49.12 s
W8-uploadsearch, product quantization50.16 s88.20 s
W12-uploadupload9.03 s30.08 s
W1 + W2upload and index, together42.2 s80.0 s1.89x

Memory and disk

What each engine held in memory and moved to and from disk over the whole run. Per-row figures are under Storage and I/O.

strawmANNQdrant
storage on disk3.1 GiB2.9 GiB
peak memory (RSS)5.2 GiB5.2 GiB
read from disk16.0 KiB1.9 GiB
written to disk2.5 GiB12.2 GiB

Peak memory and storage on disk are read before the concurrent-write rows (W11): an engine rewriting segments maps old and new files at once, and RSS counts each mapping. Bytes written are summed over the same rows, since W11's volume is set by the harness's write rate; the other disk figures are totals over every row.

Appendix

How the run was set up, every row of every table, and the diagnostics behind them. Charts on a logarithmic axis say so on the axis: read the positions there, not the distances.

Run conditions and conformance

Measured as-deployed: the scheduler was left as it ships, so a difference is what a user would see rather than the engine in isolation, and ambient load is part of the measurement. Qdrant ran equal-work: asked for one populated graph (--segments 1), as strawmANN serves, so ef means the same thing on both sides. This is the configuration §8's comparative licensing is built around, and it is a control rather than a deployment. Its optimizer treats the count as a target; what it actually held is read back from the engine and stated under Collection config.

Conformance tier reached: T4 quantization fidelity. T1 passed, so single-engine performance rows are licensed (§8). T3 passed, so a strawmANN-vs-Qdrant throughput comparison is licensed: the engines are at equal recall (§7.4). Conformance hash 0ace76ed961b7820.

On exact search the two engines' scores differ by at most 0.000e+00 (p99 0.000e+00), against a calibrated relative ε of 4.172e-07. §8.1: bit-exactness is unachievable between two different summation orders, so this — equal values within a measured tolerance — is the claim the speed rests on.

The conformance harness's single-client rate agrees on the ordering (3.54x to 3.78x); a different instrument, so no ratio is formed from it.

Every figure is the median of 3 passes per row per engine, alternated A/B/A/B (§7.4), and the spread is this run's own over 36 of 43 rows. The ingest and index-build rows (W0-upload, W1, W2, W6-upload, W7-upload, W8-upload, W12-upload) have none, so no build time carries a verdict. Each pass rebuilds the index, so the bands include build variance and are wider than a floor measured against one standing graph.

Units are queries per second. bfb reports rps, which counts batch requests: at --search-batch-size 16 the two differ by 16x, and reading one as the other once turned a 2.2x speedup into an apparent 7x regression. The tables show the request rate wherever it diverges.

The open-loop rows are not a speed. W4-sat50/70/90 offer a fixed fraction of measured saturation; serving it means the engine kept up, not that it was faster, so no ratio is printed. They exist because --parallel is a closed loop, where a stalled server stops receiving requests and understates its own tail.

ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.

What was measured builds, dataset, host
strawmANN
sm-sift-perf-0930
Qdrant
qd-sift-perf-0930
enginecommit 1fb2e437e858
ReleaseFast, native build, vnni on
binary sha256 e9f47255fbea5cee
built 2026-09-29T15:06:47Z with zig 0.16.0
version 1.19.2-dev
binary ~/.cache/strawmann/qdrant-dbeb0f73/qdrant
sha256 dbeb0f73dea2d371 build 878843e6 (from the server's banner, not a checkout)
native binary outside a checkout: sha256 is the identity (§8.9)
network native no container in the path, like strawmann
measured2026-09-29T16:00:24Z to 2026-09-29T17:44:04Z (3 passes)2026-09-29T16:20:36Z to 2026-09-29T18:07:07Z (3 passes)
profileas-deployed · both engines
load generatorbfb dev @ fc6632e5 (qdrant/bfb#176; carries #172, findings 32's --rps reaping fix) · both engines
clientqdrant-client 1.16.1-dev (git dev branch) · both engines
launched as~/Workspace/strawmann/zig-out/bin/strawmann --port 6334 --capacity 1250000 --connections 64 --workers 7 --io-threads 1 --pin --cpus 4-11 --data-dir ~/.cache/strawmann/strawmann-storage --default-placement cached~/.cache/strawmann/qdrant-dbeb0f73/qdrant
cores the engine could use4-11 (8 cores) requested by the harness
4-11 (8 cores) observed on 8 threads
0-23 (24 cores) observed on 1 thread
4-11 (8 cores) requested by the harness
4-11 (8 cores) observed on 42 threads

Dataset

sift1m, 1000000 × 128, euclid, 10000 held-out queries
ground truth: shipped, and diffed against our fp64 recompute
1 file, each pinned by sha256 in datasets.json

Host §7.1 gate pass

AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M
24 logical cores · 58.6 GB · kernel 7.0.0-34-generic
memory bandwidth: 73.5 GB/s aggregate (24 threads) · 44.8 GB/s single core (61% of bus)
strawmann startup banner, 2026-09-29T16:00:24Z; the Qdrant run agrees within 10%
strawmANN run start: environment hash 373e8531fd3d1bf6
Qdrant run start: environment hash 373e8531fd3d1bf6
All throughput rows every measurement, with its notes
workloadstrawmANNQdrantrationotes
W0d=4 floor: graph traversal with the distance taken out4,3953,5801.23x1.23x is 1.66x less work per query x 0.68x cores busy during the row (0.69 against 1.02) x 1.00x clock x 1.10x counted on-CPU share: the middle term is occupancy, not search speed d=4 floor: this measures graph traversal and heaps with the distance taken out, not request plumbing (findings 28: 76% of the server's CPU was HNSW search, 24% kernel; the RPC path is a small remainder)
W3search, fp32, single query1,7022,0080.85x
W4search, saturating (closed loop)23,05914,1351.63x
W4-sat50search, fixed rate at 50% of saturation (open loop)7,1567,156offered
W4-sat70search, fixed rate at 70% of saturation (open loop)10,01710,017offered
W4-sat90search, fixed rate at 90% of saturation (open loop)12,87912,880offered
W5search batched (16 distinct dataset queries per request)23,4001,437 requests/s4,859304 requests/s4.82xper-batch latency (16 queries/request); queries: dataset, random-sample 4.82x is 1.32x less work per query x 3.60x cores busy during the row (7.05 against 1.96) x 1.00x clock x 1.01x counted on-CPU share: the middle term is occupancy, not search speed batched 16, so qps and rps differ by 16x; the ratio uses qps
W6quantized: scalar4,3934,6900.94x
W6-ef32SQ8 recall control, ef=32 (latency only)9,3757,754-recall unequal: 0.9087 vs 0.8983; §7.4 compares at equal recall; the rescore pools match, so the encoders differ
W6-ef64SQ8 recall control, ef=64 (latency only)6,3726,179parity1.03x is 1.18x less work per query x 0.80x cores busy during the row (1.57 against 1.97) x 1.00x clock x 1.10x counted on-CPU share: the middle term is occupancy, not search speed within the ±6.1% band this dataset's noise floor puts on W6-ef64: no measured difference, not a small one
W6-ef128SQ8 recall control, ef=128 (latency only)4,4114,6840.94x
W6-ef256SQ8 recall control, ef=256 (latency only)3,0133,069paritywithin the ±2.0% band this dataset's noise floor puts on W6-ef256: no measured difference, not a small one
W6-ef512SQ8 recall control, ef=512 (latency only)1,8351,830paritywithin the ±2.0% band this dataset's noise floor puts on W6-ef512: no measured difference, not a small one
W7quantized: binary + oversampling4,5084,655-recall below the usable floor: 0.0618 and 0.0660, both under 0.50 — this row characterises the encoding, not the engines
W8quantized: PQ4,4904,1651.08x
W9exact / brute force1861241.50xbrute force over the whole collection, no index involved
W10-ef32recall control, ef=32 (latency only)40,82022,022-strawmANN: -n 100000 (2x the table's 50000) recall unequal: 0.9089 vs 0.8986; §7.4 compares at equal recall
W10-ef64recall control, ef=64 (latency only)31,42515,4942.03xstrawmANN: -n 100000 (2x the table's 50000)
W10-ef128recall control, ef=128 (latency only)20,60810,1592.03x
W10-ef256recall control, ef=256 (latency only)12,0896,5911.83x
W10-ef512recall control, ef=512 (latency only)6,6223,9631.67x
W12-sel1filtered search, one keyword (~1% of bench12)5,8073,8841.49xfiltered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set
W12-sel10filtered search, any of 10 keywords (~10% of bench12)1,3592,222-recall unequal: 1.0000 vs 0.9893; §7.4 compares at equal recall filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set
W12-sel1-ef32filtered recall control, one keyword, ef=32 (latency only)6,4196,432paritywithin the ±6.8% band this dataset's noise floor puts on W12-sel1-ef32: no measured difference, not a small one
W12-sel1-ef64filtered recall control, one keyword, ef=64 (latency only)6,2695,2581.19x
W12-sel1-ef128filtered recall control, one keyword, ef=128 (latency only)5,9963,8621.55x
W12-sel1-ef256filtered recall control, one keyword, ef=256 (latency only)6,2712,6872.33x
W12-sel1-ef512filtered recall control, one keyword, ef=512 (latency only)6,2281,7013.66x
W12-sel10-ef32filtered recall control, any of 10 keywords, ef=32 (latency only)3,1724,685-recall unequal: 0.9777 vs 0.8007; §7.4 compares at equal recall
W12-sel10-ef64filtered recall control, any of 10 keywords, ef=64 (latency only)2,0983,377-recall unequal: 0.9954 vs 0.9331; §7.4 compares at equal recall
W12-sel10-ef128filtered recall control, any of 10 keywords, ef=128 (latency only)1,3582,226-recall unequal: 1.0000 vs 0.9893; §7.4 compares at equal recall
W12-sel10-ef256filtered recall control, any of 10 keywords, ef=256 (latency only)1,3581,3251.02x
W12-sel10-ef512filtered recall control, any of 10 keywords, ef=512 (latency only)1,3597981.70x
W13scroll / pagination51,19246,5601.10xstrawmANN: -n 200000 (4x the table's 50000) Qdrant: -n 100000 (2x the table's 50000) strawmANN: 0.13 ms of the client's 0.14 ms p50 is not server time (12x), so 92% of this row is the load generator and the socket Qdrant: 0.13 ms of the client's 0.16 ms p50 is not server time (6x), so 82% of this row is the load generator and the socket 1.10x is 3.27x less work per query x 0.33x cores busy during the row (1.39 against 4.22) x 1.00x clock x 1.02x counted on-CPU share: the middle term is occupancy, not search speed
W11-steadymixed read/write below the rebuild threshold: search bench2 while 50,000 synthetic points append18,2285,537-strawmANN: qps over the 2.7 s the append ran (18,221 over the whole search); write overlap 11%; search covered 11% of the append; append 2,000 points/s Qdrant: qps over the 9.0 s the append ran (5,533 over the whole search); write overlap 36%; search covered 37% of the append; append 2,000 points/s search-during-write; no recall join; the search saw under 90% of either append, so its qps is the append's start and not all of it
W11mixed read/write: search bench2 while 200,000 synthetic points append (runs last)8,0333,800 requests/s2,387-strawmANN: qps over the 6.2 s the append ran (8,029 over the whole search); write overlap 10%; search covered 10% of the append; append 3,300 points/s Qdrant: qps over the 20.9 s the append ran (2,386 over the whole search); write overlap 35%; search covered 35% of the append; append 3,300 points/s search-during-write; no recall join; the search saw under 90% of either append, so its qps is the append's start and not all of it

W10: throughput against ef

ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.
All latency rows closed loop included, p50 to max

A row flagged “… not server time” is one where the client saw far more than the server reported: below the rate at which a queue can form, what is left is the load generator's own scheduling.

workloadstrawmANN p50strawmANN p95strawmANN p99strawmANN p99.9strawmANN maxQdrant p50Qdrant p95Qdrant p99Qdrant p99.9Qdrant max
W0d=4 floor: graph traversal with the distance taken out closed loop 223 µs251 µs272 µs331 µs890 µs272 µs324 µs352 µs414 µs1.55 ms
W3search, fp32, single query closed loop 598 µs680 µs715 µs792 µs3.01 ms505 µs564 µs594 µs666 µs2.04 ms
W4search, saturating (closed loop) closed loop 2.77 ms2.99 ms3.07 ms17.73 ms22.06 ms4.01 ms8.93 ms12.78 ms27.43 ms37.33 ms
W4-sat50search, fixed rate at 50% of saturation (open loop) open loop 493 µs834 µs1.01 ms1.56 ms2.30 ms753 µs1.38 ms1.83 ms2.32 ms5.29 ms
W4-sat70search, fixed rate at 70% of saturation (open loop) open loop 532 µs997 µs1.12 ms1.22 ms2.37 ms838 µs1.67 ms2.02 ms2.87 ms9.43 ms
W4-sat90search, fixed rate at 90% of saturation (open loop) open loop 718 µs1.26 ms1.41 ms1.53 ms3.12 ms1.24 ms3.18 ms5.27 ms9.40 ms23.12 ms
W5search batched (16 distinct dataset queries per request) closed loop 1.36 ms1.47 ms1.52 ms2.63 ms4.15 ms6.56 ms6.92 ms7.10 ms7.74 ms9.78 ms
W6quantized: scalar closed loop 442 µs573 µs605 µs649 µs1.26 ms426 µs473 µs502 µs581 µs3.04 ms
W6-ef32SQ8 recall control, ef=32 (latency only) closed loop 210 µs244 µs262 µs293 µs936 µs250 µs307 µs334 µs400 µs1.32 ms
W6-ef64SQ8 recall control, ef=64 (latency only) closed loop 315 µs377 µs401 µs440 µs1.31 ms314 µs390 µs422 µs466 µs1.47 ms
W6-ef128SQ8 recall control, ef=128 (latency only) closed loop 441 µs573 µs604 µs647 µs1.29 ms427 µs474 µs506 µs592 µs1.75 ms
W6-ef256SQ8 recall control, ef=256 (latency only) closed loop 669 µs800 µs933 µs1.01 ms1.63 ms659 µs724 µs756 µs848 µs2.04 ms
W6-ef512SQ8 recall control, ef=512 (latency only) closed loop 1.12 ms1.27 ms1.31 ms1.39 ms2.93 ms1.11 ms1.23 ms1.27 ms1.34 ms2.83 ms
W7quantized: binary + oversampling closed loop 439 µs500 µs533 µs598 µs1.13 ms427 µs485 µs513 µs584 µs1.60 ms
W8quantized: PQ closed loop 448 µs491 µs513 µs569 µs1.29 ms472 µs573 µs609 µs657 µs1.80 ms
W9exact / brute force closed loop 38.90 ms64.03 ms71.04 ms79.22 ms86.18 ms64.57 ms74.95 ms78.79 ms87.90 ms90.57 ms
W10-ef32recall control, ef=32 (latency only) closed loop 188 µs227 µs255 µs344 µs1.33 ms335 µs536 µs657 µs861 µs3.51 ms
W10-ef64recall control, ef=64 (latency only) closed loop 251 µs290 µs315 µs406 µs1.50 ms476 µs774 µs965 µs1.30 ms2.56 ms
W10-ef128recall control, ef=128 (latency only) closed loop 390 µs451 µs485 µs726 µs1.92 ms716 µs1.26 ms1.50 ms1.98 ms4.39 ms
W10-ef256recall control, ef=256 (latency only) closed loop 668 µs807 µs890 µs1.02 ms3.16 ms1.09 ms1.94 ms2.33 ms3.02 ms4.98 ms
W10-ef512recall control, ef=512 (latency only) closed loop 1.22 ms1.54 ms1.72 ms1.94 ms4.24 ms1.85 ms3.17 ms3.68 ms4.69 ms7.14 ms
W12-sel1filtered search, one keyword (~1% of bench12) closed loop 318 µs444 µs470 µs581 µs1.22 ms510 µs560 µs587 µs660 µs1.96 ms
W12-sel10filtered search, any of 10 keywords (~10% of bench12) closed loop 1.49 ms1.57 ms1.61 ms1.69 ms3.32 ms904 µs1.03 ms1.08 ms1.16 ms3.19 ms
W12-sel1-ef32filtered recall control, one keyword, ef=32 (latency only) closed loop 293 µs412 µs444 µs493 µs1.16 ms304 µs359 µs399 µs455 µs1.55 ms
W12-sel1-ef64filtered recall control, one keyword, ef=64 (latency only) closed loop 302 µs422 µs448 µs493 µs1.14 ms376 µs421 µs448 µs531 µs1.69 ms
W12-sel1-ef128filtered recall control, one keyword, ef=128 (latency only) closed loop 310 µs434 µs458 µs495 µs1.12 ms513 µs564 µs593 µs679 µs2.90 ms
W12-sel1-ef256filtered recall control, one keyword, ef=256 (latency only) closed loop 302 µs426 µs452 µs498 µs1.12 ms740 µs796 µs830 µs936 µs2.40 ms
W12-sel1-ef512filtered recall control, one keyword, ef=512 (latency only) closed loop 304 µs425 µs452 µs497 µs1.19 ms1.17 ms1.27 ms1.32 ms1.40 ms3.35 ms
W12-sel10-ef32filtered recall control, any of 10 keywords, ef=32 (latency only) closed loop 625 µs711 µs861 µs958 µs1.78 ms422 µs495 µs534 µs603 µs1.88 ms
W12-sel10-ef64filtered recall control, any of 10 keywords, ef=64 (latency only) closed loop 972 µs1.06 ms1.10 ms1.39 ms2.29 ms591 µs677 µs717 µs790 µs2.16 ms
W12-sel10-ef128filtered recall control, any of 10 keywords, ef=128 (latency only) closed loop 1.49 ms1.57 ms1.61 ms1.72 ms3.20 ms902 µs1.03 ms1.07 ms1.14 ms2.74 ms
W12-sel10-ef256filtered recall control, any of 10 keywords, ef=256 (latency only) closed loop 1.49 ms1.57 ms1.61 ms1.75 ms3.30 ms1.52 ms1.74 ms1.82 ms1.93 ms3.64 ms
W12-sel10-ef512filtered recall control, any of 10 keywords, ef=512 (latency only) closed loop 1.49 ms1.57 ms1.61 ms1.71 ms3.30 ms2.52 ms2.88 ms3.00 ms3.16 ms5.24 ms
W13scroll / pagination closed loop 137 µs204 µs230 µs270 µs1.10 ms158 µs221 µs292 µs441 µs1.52 ms
W11-steadymixed read/write below the rebuild threshold: search bench2 while 50,000 synthetic points append closed loop 394 µs571 µs1.96 ms3.01 ms6.64 ms1.26 ms2.69 ms4.79 ms9.92 ms18.90 ms
W11mixed read/write: search bench2 while 200,000 synthetic points append (runs last) closed loop 430 µs2.96 ms4.09 ms6.06 ms10.97 ms2.62 ms7.65 ms14.09 ms25.83 ms101.30 ms
Storage and I/O block layer and syscalls, per row

The syscall rows count every descriptor, sockets included, so on a search row they measure the network rather than the disk.

strawmANNQdrant
storage on disk3.1 GiBbefore W11-steady, W112.9 GiBbefore W11-steady, W11
peak RSS5.2 GiB14.7 GiB5.2 GiB before the writers
of which anonymous2.1 GiB1.7 GiB833.8 MiB before the writers
of which file-backed3.1 GiB12.8 GiB3.7 GiB before the writers
disk read bytes16.0 KiB1.9 GiB
disk write bytes2.7 GiB22.3 GiB
disk read ops1249,347
disk write ops35,822393,702
syscall reads (all fds)10,638,20846,122
syscall writes (all fds)4,883,7047,673,060

measured via strawmANN: proc+cgroup / Qdrant: proc+cgroup. proc supplies syscall counts and block-layer bytes, cgroup supplies block-layer operations and bytes, so a row one interface does not carry reads n/a via …. unknown means it was not measured, and 0 means it was: an engine started with no --data-dir has no store, which is the row this section exists for. storage on disk is the level before the rows with a concurrent writer: during those an engine that rewrites segments is caught mid-rewrite, and the same row has read 3.47 and 10.11 GiB on two runs of one binary. peak RSS does include them, being a peak.

Per workload

A search row doing block-layer reads is an engine going to disk to answer a query.

workloadstrawmANN readstrawmANN writtenstrawmANN read opsstrawmANN write opsQdrant readQdrant writtenQdrant read opsQdrant write ops
W0-upload4.0 KiB61.1 MiB3987130.2 MiB406.7 MiB3,3698,812
W000030006
W10488.4 MiB006.9 MiB1.3 GiB30921,489
W28.0 KiB488.4 MiB67,884170.8 MiB2.4 GiB4,32942,780
W300030000
W400000000
W4-sat5000000000
W4-sat7000000000
W4-sat9000000000
W500000000
W6-upload0488.4 MiB07,882161.9 MiB2.7 GiB4,10148,154
W600030000
W6-ef3200000000
W6-ef6400000000
W6-ef12800000000
W6-ef25600000000
W6-ef51200000000
W7-upload0488.4 MiB07,882194.3 MiB2.7 GiB4,91747,819
W700030000
W8-upload4.0 KiB488.4 MiB37,882145.1 MiB2.2 GiB3,62138,723
W800030000
W900000000
W10-ef3200000000
W10-ef6400000000
W10-ef12800000000
W10-ef25600000000
W10-ef51200000000
W12-upload097.8 MiB0052.0 MiB558.7 MiB1,42810,707
W12-sel10001,5770000
W12-sel1000030000
W12-sel1-ef3200000000
W12-sel1-ef6400000000
W12-sel1-ef12800000000
W12-sel1-ef25600000000
W12-sel1-ef51200000000
W12-sel10-ef3200000000
W12-sel10-ef6400000000
W12-sel10-ef12800000000
W12-sel10-ef25600000000
W12-sel10-ef51200000000
W1300000000
W11-steady024.5 MiB03365.6 MiB3.5 GiB8,97359,229
W11098.1 MiB01,707736.2 MiB6.6 GiB18,300115,983
Scheduler and memory CPU use, run-queue wait, migrations

Whether the engine was running while it ran. of wall is CPU over elapsed, waiting is runnable-but-not-scheduled, migrations checks the pinning claim, and switches gives voluntary over involuntary — the scheduler taking the core away against the engine choosing to sleep, which per unit of work is the cheapest signal of lock contention there is. A dash is not a zero: an index build's threads exit before the row does, and threads shows what fraction of the row the survivors account for.

Time spent waiting for a core

Lower is better; the bar between a pair is the gap.

Peak memory per row

Peak RSS while the row ran, from the engine's own process. Unlike the disk counters this is not refused across the two engines: residency decides where bytes live, and this is what the process held either way.

workloadstrawmANNQdrant
cpuof wallwaitingswitches vol/involmigrationsfaults min/majthreadscpuof wallwaitingswitches vol/involmigrationsfaults min/majthreads
W0-upload126.9 s566%52.7 ms47,057 / 7237114,500 / 09 1%181.7 s515%15,239.2 ms156,563 / 35,02714,207431,061 / 190,40939 5%
W08.2 s71%3.6 ms161,774 / 4600 / 0914.4 s103%3.5 ms561,464 / 3093408 / 036 0%
W12.0 s64%1.8 ms110,702 / 20160,269 / 0984.0 s604%33,583.0 ms131,555 / 50,55216,232180,060 / 259,09753 0%
W2305.7 s690%119.3 ms147,939 / 1,5751228,032 / 09 1%518.7 s626%27,187.5 ms152,557 / 42,93414,311499,998 / 298,07139 3%
W325.6 s87%7.9 ms173,310 / 4005 / 0925.7 s103%3.7 ms562,070 / 57181422 / 038 0%
W416.2 s735%1.4 ms56,793 / 28050 / 0927.5 s769%76,836.9 ms63,155 / 45,14916,5612,804 / 076
W4-sat5075.7 s270%4.4 ms292,354 / 47013 / 09122.4 s436%133,078.5 ms1,120,132 / 160,982373,210227 / 068 95%
W4-sat7070.3 s350%5.7 ms240,306 / 55018 / 09121.9 s606%167,455.3 ms1,040,817 / 213,639367,583477 / 068
W4-sat9069.8 s446%2.2 ms203,139 / 58024 / 09107.0 s683%114,892.9 ms472,846 / 122,618139,3821,409 / 044 0%
W515.1 s703%0.6 ms9,816 / 2700 / 0920.3 s197%10.2 ms33,335 / 84691393 / 042 0%
W6-upload304.7 s679%117.0 ms53,110 / 1,5551277,859 / 09 1%347.3 s545%21,501.8 ms162,119 / 38,05616,873595,971 / 298,94140 0%
W619.5 s171%3.3 ms171,686 / 3700 / 0921.6 s202%11.4 ms493,837 / 389,908551 / 044 0%
W6-ef327.5 s140%2.5 ms138,186 / 600 / 0912.8 s198%10.5 ms465,752 / 3110,249179 / 044
W6-ef6412.5 s159%10.5 ms159,938 / 1200 / 0916.1 s198%14.3 ms472,820 / 4210,36472 / 044
W6-ef12819.4 s170%4.8 ms172,966 / 1400 / 0921.7 s202%11.2 ms494,074 / 879,85495 / 044
W6-ef25629.9 s180%2.2 ms178,963 / 2400 / 0933.2 s203%10.8 ms523,021 / 12413,276148 / 044
W6-ef51251.2 s188%6.1 ms180,667 / 5900 / 0955.5 s203%15.3 ms542,027 / 19217,613364 / 044
W7-upload304.6 s679%120.6 ms53,686 / 1,6160247,539 / 09 1%352.0 s570%21,874.3 ms156,692 / 39,49118,020572,310 / 337,44243 0%
W718.7 s168%2.7 ms168,670 / 2500 / 0921.8 s203%11.5 ms503,873 / 5910,474601 / 046 0%
W8-upload376.0 s698%139.5 ms52,110 / 1,9270249,688 / 09 0%562.0 s573%10,770.0 ms2,177,353 / 80,47258,573474,960 / 240,68044 1%
W818.8 s168%8.0 ms166,890 / 2500 / 0924.3 s202%14.4 ms497,464 / 7310,385534 / 048
W974.8 s695%6.9 ms3,205 / 17500 / 09126.3 s781%8,408.8 ms17,482 / 7,9675,536220 / 050
W10-ef3211.6 s462%2.0 ms157,561 / 1200 / 0913.2 s568%5,006.8 ms228,193 / 18,12263,5721,061 / 050
W10-ef6418.9 s582%2.1 ms182,607 / 2600 / 0919.2 s588%10,527.1 ms259,766 / 28,97798,077891 / 075
W10-ef12816.3 s661%1.2 ms106,480 / 2100 / 0929.3 s591%20,877.8 ms288,104 / 35,837110,017147 / 075
W10-ef25628.9 s693%1.9 ms114,718 / 3500 / 0947.0 s617%34,985.1 ms298,809 / 37,179125,296145 / 069 0%
W10-ef51253.4 s703%4.0 ms122,860 / 8700 / 0980.3 s635%49,392.0 ms315,538 / 42,945152,309215 / 069
W12-upload48.6 s414%19.5 ms11,659 / 2790117,282 / 09 1%116.8 s339%2,585.1 ms51,033 / 5,6785,037175,816 / 152,06852 0%
W12-sel114.0 s162%8.8 ms175,682 / 18012 / 0925.9 s200%16.3 ms484,978 / 729,442499 / 057
W12-sel1069.8 s190%3.9 ms155,286 / 11800 / 0945.6 s203%12.8 ms533,768 / 15515,863396 / 057
W12-sel1-ef3212.8 s163%2.6 ms171,090 / 900 / 0915.5 s198%10.7 ms473,623 / 279,70261 / 057
W12-sel1-ef6413.0 s162%2.0 ms172,605 / 700 / 0919.1 s200%12.6 ms477,029 / 499,51765 / 057
W12-sel1-ef12813.5 s162%2.0 ms175,035 / 1100 / 0926.0 s200%17.1 ms486,023 / 879,57364 / 057
W12-sel1-ef25613.0 s162%2.5 ms175,444 / 1000 / 0937.5 s201%21.4 ms494,796 / 1269,041199 / 057
W12-sel1-ef51213.1 s162%2.4 ms175,891 / 600 / 0959.4 s202%30.5 ms521,897 / 25213,852435 / 057
W12-sel10-ef3228.1 s178%1.7 ms179,286 / 2800 / 0921.7 s202%9.0 ms507,298 / 4511,149166 / 057
W12-sel10-ef6444.2 s185%1.2 ms179,360 / 3700 / 0930.1 s203%10.2 ms519,898 / 9213,038135 / 057
W12-sel10-ef12869.8 s189%2.6 ms159,072 / 11200 / 0945.5 s202%10.5 ms534,036 / 12515,511250 / 057
W12-sel10-ef25669.8 s190%5.0 ms156,530 / 12700 / 0976.1 s202%31.4 ms546,095 / 20018,323407 / 057
W12-sel10-ef51269.8 s190%4.8 ms157,305 / 9300 / 09126.1 s201%57.6 ms554,359 / 46418,6141,068 / 057
W135.6 s138%3.2 ms243,033 / 800 / 099.1 s411%864.2 ms382,488 / 4,639112,331708 / 070
W11-steady18.9 s662%1,345.8 ms110,268 / 1340229,802 / 09 88%88.5 s968%54,590.2 ms313,517 / 88,025116,172936,740 / 224,00759 34%
W1149.5 s781%17,678.2 ms117,472 / 4,680111810,544 / 09 0%267.7 s1,271%135,269.8 ms347,677 / 138,412108,7381,842,780 / 391,25759 20%
Waiting and stalls where the time went when the engine was not running

waiting is runnable and not scheduled, blocked on disk is not runnable at all, and pressure is PSI for the engine's cgroup (some = at least one task stalled, full = every runnable task).

Time stalled on I/O

Time every task in the engine's cgroup was stalled on I/O. Lower is better. 34 rows never stalled on I/O and are not drawn.

workloadstrawmANNQdrant
waitingblocked on diskcpu some/fullio some/fullmemory some/fullwaitingblocked on diskcpu some/fullio some/fullmemory some/full
W0-upload52.7 ms0 ms15 / 10 ms1 / 1 ms0 / 0 ms15,239.2 ms0 ms1,369 / 17 ms31 / 29 ms0 / 0 ms
W03.6 ms0 ms48 / 48 ms0 / 0 ms0 / 0 ms3.5 ms0 ms80 / 80 ms0 / 0 ms0 / 0 ms
W11.8 ms0 ms11 / 11 ms0 / 0 ms0 / 0 ms33,583.0 ms0 ms2,765 / 11 ms24 / 12 ms0 / 0 ms
W2119.3 ms0 ms26 / 14 ms24 / 24 ms0 / 0 ms27,187.5 ms0 ms2,269 / 25 ms86 / 65 ms0 / 0 ms
W37.9 ms0 ms48 / 48 ms0 / 0 ms0 / 0 ms3.7 ms0 ms88 / 88 ms0 / 0 ms0 / 0 ms
W41.4 ms0 ms2 / 2 ms0 / 0 ms0 / 0 ms76,836.9 ms0 ms2,881 / 3 ms0 / 0 ms0 / 0 ms
W4-sat504.4 ms0 ms65 / 65 ms0 / 0 ms0 / 0 ms133,078.5 ms0 ms8,512 / 88 ms0 / 0 ms0 / 0 ms
W4-sat705.7 ms0 ms39 / 39 ms0 / 0 ms0 / 0 ms167,455.3 ms0 ms9,485 / 60 ms0 / 0 ms0 / 0 ms
W4-sat902.2 ms0 ms27 / 27 ms0 / 0 ms0 / 0 ms114,892.9 ms0 ms7,199 / 32 ms0 / 0 ms0 / 0 ms
W50.6 ms0 ms0 / 0 ms0 / 0 ms0 / 0 ms10.2 ms0 ms13 / 12 ms0 / 0 ms0 / 0 ms
W6-upload117.0 ms0 ms22 / 10 ms0 / 0 ms0 / 0 ms21,501.8 ms0 ms1,856 / 25 ms250 / 232 ms0 / 0 ms
W63.3 ms0 ms32 / 32 ms0 / 0 ms0 / 0 ms11.4 ms0 ms54 / 53 ms0 / 0 ms0 / 0 ms
W6-ef322.5 ms0 ms24 / 24 ms0 / 0 ms0 / 0 ms10.5 ms0 ms46 / 45 ms0 / 0 ms0 / 0 ms
W6-ef6410.5 ms0 ms30 / 30 ms0 / 0 ms0 / 0 ms14.3 ms0 ms51 / 50 ms0 / 0 ms0 / 0 ms
W6-ef1284.8 ms0 ms31 / 31 ms0 / 0 ms0 / 0 ms11.2 ms0 ms54 / 53 ms0 / 0 ms0 / 0 ms
W6-ef2562.2 ms0 ms30 / 30 ms0 / 0 ms0 / 0 ms10.8 ms0 ms59 / 59 ms0 / 0 ms0 / 0 ms
W6-ef5126.1 ms0 ms33 / 33 ms0 / 0 ms0 / 0 ms15.3 ms0 ms69 / 68 ms0 / 0 ms0 / 0 ms
W7-upload120.6 ms0 ms23 / 11 ms0 / 0 ms0 / 0 ms21,874.3 ms0 ms1,915 / 23 ms172 / 160 ms0 / 0 ms
W72.7 ms0 ms34 / 34 ms0 / 0 ms0 / 0 ms11.5 ms0 ms54 / 53 ms0 / 0 ms0 / 0 ms
W8-upload139.5 ms0 ms25 / 12 ms0 / 0 ms0 / 0 ms10,770.0 ms0 ms1,891 / 151 ms104 / 96 ms0 / 0 ms
W88.0 ms0 ms33 / 33 ms0 / 0 ms0 / 0 ms14.4 ms0 ms58 / 57 ms0 / 0 ms0 / 0 ms
W96.9 ms0 ms2 / 2 ms0 / 0 ms0 / 0 ms8,408.8 ms0 ms980 / 11 ms0 / 0 ms0 / 0 ms
W10-ef322.0 ms0 ms15 / 15 ms0 / 0 ms0 / 0 ms5,006.8 ms0 ms570 / 15 ms0 / 0 ms0 / 0 ms
W10-ef642.1 ms0 ms14 / 14 ms0 / 0 ms0 / 0 ms10,527.1 ms0 ms1,002 / 17 ms0 / 0 ms0 / 0 ms
W10-ef1281.2 ms0 ms6 / 6 ms0 / 0 ms0 / 0 ms20,877.8 ms0 ms1,789 / 20 ms0 / 0 ms0 / 0 ms
W10-ef2561.9 ms0 ms6 / 6 ms0 / 0 ms0 / 0 ms34,985.1 ms0 ms2,982 / 22 ms0 / 0 ms0 / 0 ms
W10-ef5124.0 ms0 ms6 / 6 ms0 / 0 ms0 / 0 ms49,392.0 ms0 ms4,537 / 26 ms0 / 0 ms0 / 0 ms
W12-upload19.5 ms0 ms5 / 3 ms0 / 0 ms0 / 0 ms2,585.1 ms0 ms261 / 12 ms47 / 45 ms0 / 0 ms
W12-sel18.8 ms0 ms31 / 31 ms0 / 0 ms0 / 0 ms16.3 ms0 ms56 / 54 ms0 / 0 ms0 / 0 ms
W12-sel103.9 ms0 ms36 / 36 ms0 / 0 ms0 / 0 ms12.8 ms0 ms65 / 64 ms0 / 0 ms0 / 0 ms
W12-sel1-ef322.6 ms0 ms25 / 25 ms0 / 0 ms0 / 0 ms10.7 ms0 ms49 / 49 ms0 / 0 ms0 / 0 ms
W12-sel1-ef642.0 ms0 ms27 / 27 ms0 / 0 ms0 / 0 ms12.6 ms0 ms52 / 51 ms0 / 0 ms0 / 0 ms
W12-sel1-ef1282.0 ms0 ms29 / 29 ms0 / 0 ms0 / 0 ms17.1 ms0 ms56 / 54 ms0 / 0 ms0 / 0 ms
W12-sel1-ef2562.5 ms0 ms27 / 27 ms0 / 0 ms0 / 0 ms21.4 ms0 ms60 / 58 ms0 / 0 ms0 / 0 ms
W12-sel1-ef5122.4 ms0 ms27 / 27 ms0 / 0 ms0 / 0 ms30.5 ms0 ms72 / 69 ms0 / 0 ms0 / 0 ms
W12-sel10-ef321.7 ms0 ms31 / 31 ms0 / 0 ms0 / 0 ms9.0 ms0 ms54 / 53 ms0 / 0 ms0 / 0 ms
W12-sel10-ef641.2 ms0 ms31 / 31 ms0 / 0 ms0 / 0 ms10.2 ms0 ms58 / 57 ms0 / 0 ms0 / 0 ms
W12-sel10-ef1282.6 ms0 ms36 / 36 ms0 / 0 ms0 / 0 ms10.5 ms0 ms64 / 63 ms0 / 0 ms0 / 0 ms
W12-sel10-ef2565.0 ms0 ms36 / 36 ms0 / 0 ms0 / 0 ms31.4 ms0 ms95 / 90 ms0 / 0 ms0 / 0 ms
W12-sel10-ef5124.8 ms0 ms36 / 36 ms0 / 0 ms0 / 0 ms57.6 ms0 ms124 / 117 ms0 / 0 ms0 / 0 ms
W133.2 ms0 ms28 / 28 ms0 / 0 ms0 / 0 ms864.2 ms0 ms121 / 26 ms0 / 0 ms0 / 0 ms
W11-steady1,345.8 ms0 ms258 / 13 ms0 / 0 ms0 / 0 ms54,590.2 ms0 ms4,204 / 27 ms212 / 193 ms0 / 0 ms
W1117,678.2 ms0 ms3,366 / 25 ms10 / 10 ms0 / 0 ms135,269.8 ms0 ms10,206 / 61 ms412 / 318 ms2 / 1 ms

A dash under blocked on disk is not a zero: most hosts ship with kernel.task_delayacct=0, which reports the field as a permanent zero and would manufacture the strongest claim here — that the memory-resident engine never waits for a device — out of a sysctl. bench/setup.py apply turns it on.

Hardware counters cycles, DRAM loads, TLB walks, IPC per query

What the core did per query, from perf stat attached to the engine for the same window as the /proc counters. §5's cost model is a set of claims about instructions per cycle, DRAM traffic and TLB reach; these are those quantities, measured on the row the headline quotes rather than on a microbenchmark.

What one query cost, in cycles

Lower is better; the bar between a pair is the gap. Load rows are absent: they have no queries to divide by, and an absolute count under a per-query axis would be a different quantity wearing this one's label.

Demand loads from DRAM per query

Demand loads served from DRAM, times the cache line this host reports. Lower is better. Demand only: the hardware prefetcher's fills are a separate counter and are not in this number, so a row that streams — an exact scan above all — moved far more than this says. On the graph rows, where there is little for a prefetcher to predict, it is most of the traffic and is the outstanding-miss quantity §5.2 argues the engine is limited by.

TLB walks per query

Data TLB misses that reached a page walk, per query. Lower is better. §5.5 argues a memory-resident index needs hugepage care; this is what not taking it costs, and `--no-huge-pages` is the A/B arm that prices it.

Instructions per cycle

How well the core was fed while it ran. Not a score: a low IPC on a memory-bound row is what §5.2 predicts, and a high one on a row that does more work per query is not a win. Read it beside the two charts above.

Branch mispredictions per 1k instructions

A graph traversal is a chain of data-dependent branches and none of them is predictable from the last query, so this is the column that separates a slow kernel from an unpredictable walk. Per thousand instructions rather than per branch: `branches` costs a PMU counter that `ref-cycles` needs more, and MPKI is the comparable form anyway.

workloadstrawmANNQdrant
IPCGHzbranch MPKIcycles/querydemand DRAM/queryTLB walks/queryIPCGHzbranch MPKIcycles/querydemand DRAM/queryTLB walks/query
W0-upload1.522.011.005x nominal5.40---1.332.011.005x nominal4.69---
W01.272.011.005x nominal6.06273,3807.4 KiB1,213.01.262.011.005x nominal3.22452,84746.2 KiB1,319.0
W11.802.011.005x nominal0.80---1.412.011.005x nominal3.94---
W21.292.001.004x nominal4.53---1.082.011.004x nominal4.02---
W30.752.011.005x nominal6.06947,761776.9 KiB3,340.21.192.011.005x nominal3.20900,7591,023.0 KiB3,671.2
W41.082.011.006x nominal5.25641,138765.9 KiB2,816.90.912.011.006x nominal3.741,045,2751,104.6 KiB3,102.1
W4-sat500.942.011.006x nominal5.94739,002777.9 KiB3,077.90.962.011.006x nominal3.791,093,9041,100.1 KiB3,539.4
W4-sat701.012.011.005x nominal5.71686,638775.1 KiB2,980.10.962.011.005x nominal3.761,088,2571,102.4 KiB3,420.5
W4-sat901.012.011.005x nominal5.71685,960774.2 KiB2,932.10.962.011.005x nominal3.721,012,5691,076.0 KiB3,213.4
W51.062.001.004x nominal5.59602,015776.1 KiB2,827.01.002.011.006x nominal4.33795,6681,088.9 KiB2,787.6
W6-upload1.292.011.005x nominal4.51---1.802.011.005x nominal4.56---
W61.002.011.005x nominal6.59727,296172.5 KiB4,507.31.702.011.006x nominal3.08754,218126.0 KiB3,176.8
W6-ef321.042.011.006x nominal3.61270,63850.2 KiB1,791.61.602.011.006x nominal2.24405,96639.7 KiB1,458.2
W6-ef640.952.011.006x nominal5.28457,56693.4 KiB2,799.31.602.011.006x nominal2.64540,47168.1 KiB2,148.9
W6-ef1281.002.011.006x nominal6.59724,184172.2 KiB4,508.41.692.011.006x nominal3.12753,540127.5 KiB3,176.6
W6-ef2561.122.011.005x nominal7.201,143,584310.7 KiB7,747.71.692.011.006x nominal3.591,207,544236.6 KiB4,938.1
W6-ef5121.182.011.005x nominal7.681,986,580566.3 KiB14,644.61.642.011.006x nominal4.012,090,557453.1 KiB8,019.3
W7-upload1.292.001.004x nominal4.55---2.122.011.005x nominal3.95---
W71.282.011.006x nominal7.65696,49254.6 KiB4,029.01.772.011.006x nominal3.30762,270103.5 KiB2,896.8
W8-upload2.392.001.004x nominal2.23---4.502.001.004x nominal0.93---
W82.142.011.006x nominal3.45701,32456.7 KiB2,820.22.182.011.006x nominal1.97857,079102.8 KiB2,838.3
W92.012.011.005x nominal0.0174,904,81634,264.3 KiB33,969.31.332.011.006x nominal0.15125,896,18110,931.2 KiB126,608.9
W10-ef321.252.011.006x nominal3.60217,779243.1 KiB978.71.142.011.005x nominal2.66458,682363.5 KiB1,277.9
W10-ef641.172.011.006x nominal4.45361,769424.7 KiB1,611.61.042.011.006x nominal3.19682,559630.7 KiB2,049.3
W10-ef1281.112.011.005x nominal5.23632,461767.1 KiB2,860.70.972.011.006x nominal3.751,065,7671,094.2 KiB3,362.0
W10-ef2561.072.011.006x nominal5.941,138,5221,396.4 KiB5,391.80.932.011.005x nominal4.271,752,4451,913.4 KiB5,486.9
W10-ef5121.032.011.006x nominal6.572,116,5742,527.4 KiB10,566.60.892.011.005x nominal4.723,063,2313,399.6 KiB9,095.5
W12-upload1.412.011.005x nominal5.00---1.592.011.005x nominal3.90---
W12-sel11.032.011.006x nominal4.88511,298693.8 KiB2,279.31.332.011.006x nominal4.03925,913411.5 KiB2,568.2
W12-sel101.312.011.006x nominal1.142,738,3087,811.4 KiB7,120.11.422.011.006x nominal2.731,700,049924.7 KiB3,907.2
W12-sel1-ef321.122.011.006x nominal4.60469,316681.0 KiB2,232.21.292.011.006x nominal2.98513,455208.6 KiB1,577.6
W12-sel1-ef641.102.011.006x nominal4.69478,534687.3 KiB2,249.21.322.011.006x nominal3.53655,811301.7 KiB2,055.7
W12-sel1-ef1281.052.011.006x nominal4.76496,626690.6 KiB2,253.61.322.011.006x nominal4.04931,805420.3 KiB2,563.9
W12-sel1-ef2561.112.011.006x nominal4.68474,473685.9 KiB2,261.31.382.011.006x nominal4.461,385,109548.5 KiB3,120.4
W12-sel1-ef5121.102.011.005x nominal4.74477,695687.3 KiB2,267.21.412.011.006x nominal4.832,250,713741.8 KiB3,744.4
W12-sel10-ef321.512.011.005x nominal3.571,066,755385.2 KiB2,742.41.412.011.006x nominal2.40752,520339.6 KiB2,060.3
W12-sel10-ef641.482.011.006x nominal4.051,710,408669.7 KiB3,903.11.412.011.006x nominal2.541,084,807555.5 KiB2,791.9
W12-sel10-ef1281.312.011.006x nominal1.142,738,1937,806.5 KiB7,120.01.422.011.006x nominal2.721,698,423923.2 KiB3,909.8
W12-sel10-ef2561.312.011.006x nominal1.142,738,6077,812.9 KiB7,117.41.392.011.006x nominal3.142,898,5691,545.5 KiB5,491.6
W12-sel10-ef5121.312.011.006x nominal1.142,737,5267,808.8 KiB7,128.41.432.011.005x nominal3.314,866,3862,421.6 KiB7,610.4
W131.362.011.007x nominal1.2242,2341.0 KiB13.11.472.011.006x nominal1.40137,9991.4 KiB41.8
W11-steady1.232.011.005x nominal4.74732,549802.7 KiB2,974.41.912.011.005x nominal2.013,380,2701,691.8 KiB4,485.8
W111.722.011.004x nominal2.491,913,3921,614.5 KiB3,878.71.792.011.004x nominal1.6410,457,6056,698.1 KiB17,066.6

Nothing here is scaled. Where perf had to multiplex the group, the values it prints are extrapolations from the fraction of the row each counter was on, and they are withheld rather than shown — the same rule the scheduler table applies to a thread that exited. An event this host does not implement is likewise blank, never zero.

The two instruments disagree on these rows. strawmANN: perf and /proc disagree on context switches by up to 16% on W11; Qdrant: perf and /proc disagree on minor faults by up to 10% on W2, W6-upload, W7-upload, W8-upload, W12-upload, W11-steady, W11; Qdrant: perf and /proc disagree on context switches by up to 13% on W8-upload. perf keeps an exiting thread's counts and the /proc sums do not, so a gap here is usually the same lost-thread effect the scheduler table reports as coverage.

Collection config what each engine says it built

Read back from the engine rather than taken from the request. --segments 1 sets Qdrant's default_segment_number, which its optimizer treats as a target, and the segment count is the largest confound in the ef comparison. strawmANN held 1 segment; Qdrant held 2 segments.

Where the engines disagree a cell reads strawmANN / Qdrant, and is marked.

collectionsegmentspopulated segmentsrequested segmentspointsindexed vectorsvector sizehnsw mhnsw ef_constructquantizationshardsvector residency
bench01 / 2- / 1- / 11,000,0001,000,000416100-1cached
bench11 / 5- / 5- / 11,000,0000 / 101,10012816100-1cached
bench21 / 2- / 1- / 11,000,0001,000,00012816100-1cached
bench61 / 2- / 1- / 11,000,0001,000,00012816100scalar1cached
bench71 / 2- / 1- / 11,000,0001,000,00012816100binary1cached
bench81 / 2- / 1- / 11,000,0001,000,00012816100product1cached
bench121 / 2- / 1- / 1200,000200,00012816100-1cached

After the mutating rows bench2 held strawmANN 1,250,000 (+250,000), 1,070,900 of them indexed; Qdrant 1,250,000 (+250,000), 1,223,700 of them indexed. The table above is the state the throughput, latency and recall rows searched; this is what W11 left, and is that row's subject rather than theirs.

bench1 was read back straight after its row and dropped, with no settle between, so an engine still building it reports a snapshot taken mid-ingest. Its cells are not marked as a disagreement.

Per workload every metric of every row, engine beside engine

Metric down, engine across, so a comparison is two adjacent cells. Metrics a row did not measure are dropped rather than shown empty.

W0-upload transport floor: load

upload · collection=bench0 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=1,000,000
metricstrawmANNQdrant
wall clock22.4 s35.3 s1.57x
of which upload1.3 s7.2 s5.59x
of which index wait19.1 s26.1 s1.37x
time to Green19.1 s26.1 s1.37x
cpu126.9 s181.7 s1.43x
cpu, % of wall566%515%
waiting for a core52.7 ms15,239.2 ms289.41x
migrations714,2072,029.57x
faults min/maj114,500 / 0431,061 / 190,409
peak RSS443.8 MiB1.5 GiB3.37x
disk written61.1 MiB406.7 MiB
disk read4.0 KiB130.2 MiB

W0 d=4 floor: graph traversal with the distance taken out 1.23x

closed-loop · collection=bench0 · client -p 1 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second4,3953,5801.23x
client p50223 µs272 µs1.22x
client p99272 µs352 µs1.29x
client p99.9331 µs414 µs1.25x
server p50132 µs171 µs1.30x
server p99145 µs228 µs1.57x
wall clock11.4 s14.0 s1.23x
queries sent50,00050,000
cpu8.2 s14.4 s1.76x
cpu, % of wall71%103%
waiting for a core3.6 ms3.5 ms0.95x
migrations093
faults min/maj0 / 0408 / 0
peak RSS443.8 MiB1.5 GiB3.37x
disk written00
disk read00
1.23x is 1.66x less work per query x 0.68x cores busy during the row (0.69 against 1.02) x 1.00x clock x 1.10x counted on-CPU share: the middle term is occupancy, not search speed d=4 floor: this measures graph traversal and heaps with the distance taken out, not request plumbing (findings 28: 76% of the server's CPU was HNSW search, 24% kernel; the RPC path is a small remainder)

W1 ingest throughput

upload · collection=bench1 · client -p 8 -t 8 -c 1 (-c bfb default) · n=1,000,000
metricstrawmANNQdrant
wall clock3.2 s13.9 s4.34x
of which upload1.1 s11.8 s10.96x
cpu2.0 s84.0 s41.20x
cpu, % of wall64%604%
waiting for a core1.8 ms33,583.0 ms18,323.08x
migrations016,232
faults min/maj160,269 / 0180,060 / 259,097
peak RSS1.1 GiB2.3 GiB2.04x
disk written488.4 MiB1.3 GiB
disk read06.9 MiB

W2 index build time

upload · collection=bench2 · client -p 8 -t 8 -c 1 (-c bfb default) · n=1,000,000
metricstrawmANNQdrant
wall clock44.3 s82.8 s1.87x
of which upload1.1 s12.3 s11.27x
of which index wait41.1 s68.2 s1.66x
time to Green41.1 s68.2 s1.66x
cpu305.7 s518.7 s1.70x
cpu, % of wall690%626%
waiting for a core119.3 ms27,187.5 ms227.95x
migrations114,31114,311.00x
faults min/maj228,032 / 0499,998 / 298,071
peak RSS1.4 GiB3.2 GiB2.36x
disk written488.4 MiB2.4 GiB
disk read8.0 KiB170.8 MiB

W3 search, fp32, single query 0.85x

ef=128 · closed-loop · collection=bench2 · client -p 1 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,7022,0080.85x
recall@100.98880.9875~equal
client p50598 µs505 µs0.84x
client p99715 µs594 µs0.83x
client p99.9792 µs666 µs0.84x
server p50492 µs409 µs0.83x
server p99597 µs481 µs0.81x
wall clock29.4 s24.9 s0.85x
queries sent50,00050,000
cpu25.6 s25.7 s~equal
cpu, % of wall87%103%
waiting for a core7.9 ms3.7 ms0.46x
migrations0181
faults min/maj5 / 0422 / 0
peak RSS1.4 GiB3.2 GiB2.36x
disk written00
disk read00

W4 search, saturating (closed loop) 1.63x

ef=128 · closed-loop · collection=bench2 · client -p 64 -t 16 -c 2 · n=50,000
metricstrawmANNQdrant
queries/second23,05914,1351.63x
recall@100.98880.9875~equal
client p502.77 ms4.01 ms1.45x
client p993.07 ms12.78 ms4.16x
client p99.917.73 ms27.43 ms1.55x
server p502.69 ms2.80 ms1.04x
server p992.97 ms9.77 ms3.28x
wall clock2.2 s3.6 s1.62x
queries sent50,00050,000
cpu16.2 s27.5 s1.70x
cpu, % of wall735%769%
waiting for a core1.4 ms76,836.9 ms55,996.19x
migrations016,561
faults min/maj50 / 02,804 / 0
peak RSS1.4 GiB3.2 GiB2.36x
disk written00
disk read00

W4-sat50 search, fixed rate at 50% of saturation (open loop) offered

ef=128 · open-loop · offered=7,156/s (50% of saturation, pinned reference 14,312 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second7,1567,156offered
recall@100.98880.9875~equal
client p50493 µs753 µs1.53x
client p991.01 ms1.83 ms1.82x
client p99.91.56 ms2.32 ms1.48x
server p50380 µs588 µs1.55x
server p99852 µs1.57 ms1.84x
wall clock28.1 s28.1 s~equal
queries sent200,000200,000
cpu75.7 s122.4 s1.62x
cpu, % of wall270%436%
waiting for a core4.4 ms133,078.5 ms30,255.65x
migrations0373,210
faults min/maj13 / 0227 / 0
peak RSS1.4 GiB3.2 GiB2.36x
disk written00
disk read00

W4-sat70 search, fixed rate at 70% of saturation (open loop) offered

ef=128 · open-loop · offered=10,018/s (70% of saturation, pinned reference 14,312 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second10,01710,017offered
recall@100.98880.9875~equal
client p50532 µs838 µs1.57x
client p991.12 ms2.02 ms1.80x
client p99.91.22 ms2.87 ms2.35x
server p50410 µs641 µs1.56x
server p99942 µs1.72 ms1.83x
wall clock20.1 s20.1 s~equal
queries sent200,000200,000
cpu70.3 s121.9 s1.73x
cpu, % of wall350%606%
waiting for a core5.7 ms167,455.3 ms29,377.87x
migrations0367,583
faults min/maj18 / 0477 / 0
peak RSS1.4 GiB3.2 GiB2.36x
disk written00
disk read00

W4-sat90 search, fixed rate at 90% of saturation (open loop) offered

ef=128 · open-loop · offered=12,881/s (90% of saturation, pinned reference 14,312 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second12,87912,880offered
recall@100.98880.9875~equal
client p50718 µs1.24 ms1.72x
client p991.41 ms5.27 ms3.73x
client p99.91.53 ms9.40 ms6.13x
server p50597 µs887 µs1.49x
server p991.22 ms3.80 ms3.11x
wall clock15.7 s15.7 s~equal
queries sent200,000200,000
cpu69.8 s107.0 s1.53x
cpu, % of wall446%683%
waiting for a core2.2 ms114,892.9 ms52,563.19x
migrations0139,382
faults min/maj24 / 01,409 / 0
peak RSS1.4 GiB3.2 GiB2.36x
disk written00
disk read00

W5 search batched (16 distinct dataset queries per request) 4.82x

ef=128 · closed-loop · collection=bench2 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second23,4004,8594.82x
recall@100.98880.9875~equal
client p501.36 ms6.56 ms4.83x
client p991.52 ms7.10 ms4.67x
client p99.92.63 ms7.74 ms2.94x
server p501.22 ms6.17 ms5.07x
server p991.38 ms6.61 ms4.80x
wall clock2.1 s10.3 s4.80x
queries sent50,00050,000
cpu15.1 s20.3 s1.35x
cpu, % of wall703%197%
waiting for a core0.6 ms10.2 ms15.73x
migrations0691
faults min/maj0 / 0393 / 0
peak RSS1.4 GiB3.2 GiB2.36x
disk written00
disk read00
per-batch latency (16 queries/request); queries: dataset, random-sample 4.82x is 1.32x less work per query x 3.60x cores busy during the row (7.05 against 1.96) x 1.00x clock x 1.01x counted on-CPU share: the middle term is occupancy, not search speed batched 16, so qps and rps differ by 16x; the ratio uses qps

W6-upload scalar quantization: load

upload · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=1,000,000
metricstrawmANNQdrant
wall clock44.9 s63.7 s1.42x
of which upload1.6 s10.1 s6.14x
of which index wait41.1 s53.1 s1.29x
time to Green41.1 s53.1 s1.29x
cpu304.7 s347.3 s1.14x
cpu, % of wall679%545%
waiting for a core117.0 ms21,501.8 ms183.73x
migrations116,87316,873.00x
faults min/maj277,859 / 0595,971 / 298,941
peak RSS2.4 GiB4.6 GiB1.88x
disk written488.4 MiB2.7 GiB
disk read0161.9 MiB

W6 quantized: scalar 0.94x

ef=128 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second4,3934,6900.94x
recall@100.98900.9871~equal
client p50442 µs426 µs0.96x
client p99605 µs502 µs0.83x
client p99.9649 µs581 µs0.89x
server p50349 µs329 µs0.94x
server p99495 µs386 µs0.78x
wall clock11.4 s10.7 s0.94x
queries sent50,00050,000
cpu19.5 s21.6 s1.11x
cpu, % of wall171%202%
waiting for a core3.3 ms11.4 ms3.47x
migrations09,908
faults min/maj0 / 0551 / 0
peak RSS2.4 GiB4.6 GiB1.88x
disk written00
disk read00

W6-ef32 SQ8 recall control, ef=32 (latency only)

ef=32 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=3.2 · rescore=true
metricstrawmANNQdrant
queries/second9,3757,754-
recall@100.90870.8983
client p50210 µs250 µs
client p99262 µs334 µs
client p99.9293 µs400 µs
server p50117 µs153 µs
server p99156 µs213 µs
wall clock5.4 s6.5 s
queries sent50,00050,000
cpu7.5 s12.8 s
cpu, % of wall140%198%
waiting for a core2.5 ms10.5 ms
migrations010,249
faults min/maj0 / 0179 / 0
peak RSS2.4 GiB4.6 GiB
disk written00
disk read00
recall unequal: 0.9087 vs 0.8983; §7.4 compares at equal recall; the rescore pools match, so the encoders differ

W6-ef64 SQ8 recall control, ef=64 (latency only) parity

ef=64 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=6.4 · rescore=true
metricstrawmANNQdrant
queries/second6,3726,179parity
recall@100.96460.9600~equal
client p50315 µs314 µs~equal
client p99401 µs422 µs1.05x
client p99.9440 µs466 µs1.06x
server p50217 µs216 µs~equal
server p99291 µs299 µs1.03x
wall clock7.9 s8.1 s1.03x
queries sent50,00050,000
cpu12.5 s16.1 s1.28x
cpu, % of wall159%198%
waiting for a core10.5 ms14.3 ms1.36x
migrations010,364
faults min/maj0 / 072 / 0
peak RSS2.4 GiB4.6 GiB1.88x
disk written00
disk read00
1.03x is 1.18x less work per query x 0.80x cores busy during the row (1.57 against 1.97) x 1.00x clock x 1.10x counted on-CPU share: the middle term is occupancy, not search speed within the ±6.1% band this dataset's noise floor puts on W6-ef64: no measured difference, not a small one

W6-ef128 SQ8 recall control, ef=128 (latency only) 0.94x

ef=128 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second4,4114,6840.94x
recall@100.98900.9871~equal
client p50441 µs427 µs0.97x
client p99604 µs506 µs0.84x
client p99.9647 µs592 µs0.91x
server p50348 µs329 µs0.95x
server p99494 µs386 µs0.78x
wall clock11.4 s10.7 s0.94x
queries sent50,00050,000
cpu19.4 s21.7 s1.12x
cpu, % of wall170%202%
waiting for a core4.8 ms11.2 ms2.32x
migrations09,854
faults min/maj0 / 095 / 0
peak RSS2.4 GiB4.6 GiB1.88x
disk written00
disk read00

W6-ef256 SQ8 recall control, ef=256 (latency only) parity

ef=256 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=25.6 · rescore=true
metricstrawmANNQdrant
queries/second3,0133,069parity
recall@100.99740.9968~equal
client p50669 µs659 µs~equal
client p99933 µs756 µs0.81x
client p99.91.01 ms848 µs0.84x
server p50578 µs561 µs0.97x
server p99825 µs641 µs0.78x
wall clock16.6 s16.3 s~equal
queries sent50,00050,000
cpu29.9 s33.2 s1.11x
cpu, % of wall180%203%
waiting for a core2.2 ms10.8 ms4.92x
migrations013,276
faults min/maj0 / 0148 / 0
peak RSS2.4 GiB4.6 GiB1.88x
disk written00
disk read00
within the ±2.0% band this dataset's noise floor puts on W6-ef256: no measured difference, not a small one

W6-ef512 SQ8 recall control, ef=512 (latency only) parity

ef=512 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=51.2 · rescore=true
metricstrawmANNQdrant
queries/second1,8351,830parity
recall@100.99960.9993~equal
client p501.12 ms1.11 ms~equal
client p991.31 ms1.27 ms0.96x
client p99.91.39 ms1.34 ms0.96x
server p501.02 ms1.01 ms~equal
server p991.22 ms1.15 ms0.95x
wall clock27.3 s27.4 s~equal
queries sent50,00050,000
cpu51.2 s55.5 s1.08x
cpu, % of wall188%203%
waiting for a core6.1 ms15.3 ms2.52x
migrations017,613
faults min/maj0 / 0364 / 0
peak RSS2.4 GiB4.6 GiB1.88x
disk written00
disk read00
within the ±2.0% band this dataset's noise floor puts on W6-ef512: no measured difference, not a small one

W7-upload binary quantization: load

upload · collection=bench7 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=1,000,000
metricstrawmANNQdrant
wall clock44.9 s61.7 s1.38x
of which upload1.6 s10.5 s6.40x
of which index wait41.1 s49.1 s1.19x
time to Green41.1 s49.1 s1.19x
cpu304.6 s352.0 s1.16x
cpu, % of wall679%570%
waiting for a core120.6 ms21,874.3 ms181.45x
migrations018,020
faults min/maj247,539 / 0572,310 / 337,442
peak RSS3.4 GiB5.2 GiB1.54x
disk written488.4 MiB2.7 GiB
disk read0194.3 MiB

W7 quantized: binary + oversampling

ef=128 · closed-loop · collection=bench7 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second4,5084,655-
recall@100.06180.0660
client p50439 µs427 µs
client p99533 µs513 µs
client p99.9598 µs584 µs
server p50337 µs329 µs
server p99424 µs403 µs
wall clock11.1 s10.8 s
queries sent50,00050,000
cpu18.7 s21.8 s
cpu, % of wall168%203%
waiting for a core2.7 ms11.5 ms
migrations010,474
faults min/maj0 / 0601 / 0
peak RSS3.4 GiB5.2 GiB
disk written00
disk read00
recall below the usable floor: 0.0618 and 0.0660, both under 0.50 — this row characterises the encoding, not the engines

W8-upload PQ: load

upload · collection=bench8 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=1,000,000
metricstrawmANNQdrant
wall clock53.9 s98.1 s1.82x
of which upload1.6 s7.8 s4.83x
of which index wait50.2 s88.2 s1.76x
time to Green50.2 s88.2 s1.76x
cpu376.0 s562.0 s1.49x
cpu, % of wall698%573%
waiting for a core139.5 ms10,770.0 ms77.21x
migrations058,573
faults min/maj249,688 / 0474,960 / 240,680
peak RSS4.3 GiB5.2 GiB1.20x
disk written488.4 MiB2.2 GiB
disk read4.0 KiB145.1 MiB

W8 quantized: PQ 1.08x

ef=128 · closed-loop · collection=bench8 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second4,4904,1651.08x
recall@100.98250.9776~equal
client p50448 µs472 µs1.05x
client p99513 µs609 µs1.19x
client p99.9569 µs657 µs1.16x
server p50348 µs373 µs1.07x
server p99396 µs486 µs1.23x
wall clock11.2 s12.0 s1.08x
queries sent50,00050,000
cpu18.8 s24.3 s1.29x
cpu, % of wall168%202%
waiting for a core8.0 ms14.4 ms1.80x
migrations010,385
faults min/maj0 / 0534 / 0
peak RSS4.3 GiB5.2 GiB1.20x
disk written00
disk read00

W9 exact / brute force 1.50x

exact · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=2,000
metricstrawmANNQdrant
queries/second1861241.50x
client p5038.90 ms64.57 ms1.66x
client p9971.04 ms78.79 ms1.11x
client p99.979.22 ms87.90 ms1.11x
server p5038.31 ms63.45 ms1.66x
server p9970.31 ms77.79 ms1.11x
wall clock10.8 s16.2 s1.50x
queries sent2,0002,000
cpu74.8 s126.3 s1.69x
cpu, % of wall695%781%
waiting for a core6.9 ms8,408.8 ms1,211.84x
migrations05,536
faults min/maj0 / 0220 / 0
peak RSS4.3 GiB5.2 GiB1.20x
disk written00
disk read00
brute force over the whole collection, no index involved

W10-ef32 recall control, ef=32 (latency only)

ef=32 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=100,000
metricstrawmANNQdrant
queries/second40,82022,022-
recall@100.90890.8986
client p50188 µs335 µs
client p99255 µs657 µs
client p99.9344 µs861 µs
server p5096 µs208 µs
server p99132 µs447 µs
wall clock2.5 s2.3 s
queries sent100,00050,000
cpu11.6 s13.2 s
cpu, % of wall462%568%
waiting for a core2.0 ms5,006.8 ms
migrations063,572
faults min/maj0 / 01,061 / 0
peak RSS4.3 GiB5.2 GiB
disk written00
disk read00
strawmANN: -n 100000 (2x the table's 50000) recall unequal: 0.9089 vs 0.8986; §7.4 compares at equal recall

W10-ef64 recall control, ef=64 (latency only) 2.03x

ef=64 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=100,000
metricstrawmANNQdrant
queries/second31,42515,4942.03x
recall@100.96390.9603~equal
client p50251 µs476 µs1.90x
client p99315 µs965 µs3.07x
client p99.9406 µs1.30 ms3.20x
server p50170 µs335 µs1.97x
server p99215 µs761 µs3.55x
wall clock3.2 s3.3 s~equal
queries sent100,00050,000
cpu18.9 s19.2 s~equal
cpu, % of wall582%588%
waiting for a core2.1 ms10,527.1 ms4,907.24x
migrations098,077
faults min/maj0 / 0891 / 0
peak RSS4.3 GiB5.2 GiB1.20x
disk written00
disk read00
strawmANN: -n 100000 (2x the table's 50000)

W10-ef128 recall control, ef=128 (latency only) 2.03x

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second20,60810,1592.03x
recall@100.98880.9875~equal
client p50390 µs716 µs1.83x
client p99485 µs1.50 ms3.09x
client p99.9726 µs1.98 ms2.73x
server p50313 µs561 µs1.79x
server p99401 µs1.28 ms3.19x
wall clock2.5 s5.0 s2.01x
queries sent50,00050,000
cpu16.3 s29.3 s1.80x
cpu, % of wall661%591%
waiting for a core1.2 ms20,877.8 ms17,186.58x
migrations0110,017
faults min/maj0 / 0147 / 0
peak RSS4.3 GiB5.2 GiB1.20x
disk written00
disk read00

W10-ef256 recall control, ef=256 (latency only) 1.83x

ef=256 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second12,0896,5911.83x
recall@100.99740.9967~equal
client p50668 µs1.09 ms1.63x
client p99890 µs2.33 ms2.62x
client p99.91.02 ms3.02 ms2.95x
server p50587 µs923 µs1.57x
server p99809 µs2.03 ms2.51x
wall clock4.2 s7.6 s1.83x
queries sent50,00050,000
cpu28.9 s47.0 s1.63x
cpu, % of wall693%617%
waiting for a core1.9 ms34,985.1 ms18,362.54x
migrations0125,296
faults min/maj0 / 0145 / 0
peak RSS4.3 GiB5.2 GiB1.20x
disk written00
disk read00

W10-ef512 recall control, ef=512 (latency only) 1.67x

ef=512 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second6,6223,9631.67x
recall@100.99950.9992~equal
client p501.22 ms1.85 ms1.52x
client p991.72 ms3.68 ms2.14x
client p99.91.94 ms4.69 ms2.42x
server p501.13 ms1.65 ms1.46x
server p991.63 ms3.35 ms2.05x
wall clock7.6 s12.7 s1.67x
queries sent50,00050,000
cpu53.4 s80.3 s1.51x
cpu, % of wall703%635%
waiting for a core4.0 ms49,392.0 ms12,311.47x
migrations0152,309
faults min/maj0 / 0215 / 0
peak RSS4.3 GiB5.2 GiB1.20x
disk written00
disk read00

W12-upload filtered search: load 200,000 with payloads

upload · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=200,000
metricstrawmANNQdrant
wall clock11.7 s34.4 s2.93x
of which upload0.6 s2.0 s3.28x
of which index wait9.0 s30.1 s3.33x
time to Green9.0 s30.1 s3.33x
cpu48.6 s116.8 s2.40x
cpu, % of wall414%339%
waiting for a core19.5 ms2,585.1 ms132.37x
migrations05,037
faults min/maj117,282 / 0175,816 / 152,068
peak RSS5.2 GiB5.2 GiB~equal
disk written97.8 MiB558.7 MiB
disk read052.0 MiB

W12-sel1 filtered search, one keyword (~1% of bench12) 1.49x

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second5,8073,8841.49x
recall@101.00001.0000~equal
client p50318 µs510 µs1.61x
client p99470 µs587 µs1.25x
client p99.9581 µs660 µs1.14x
server p50228 µs411 µs1.80x
server p99352 µs470 µs1.33x
wall clock8.6 s12.9 s1.49x
queries sent50,00050,000
cpu14.0 s25.9 s1.84x
cpu, % of wall162%200%
waiting for a core8.8 ms16.3 ms1.85x
migrations09,442
faults min/maj12 / 0499 / 0
peak RSS5.2 GiB5.2 GiB~equal
disk written00
disk read00
filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set

W12-sel10 filtered search, any of 10 keywords (~10% of bench12)

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,3592,222-
recall@101.00000.9893
client p501.49 ms904 µs
client p991.61 ms1.08 ms
client p99.91.69 ms1.16 ms
server p501.39 ms796 µs
server p991.50 ms963 µs
wall clock36.8 s22.5 s
queries sent50,00050,000
cpu69.8 s45.6 s
cpu, % of wall190%203%
waiting for a core3.9 ms12.8 ms
migrations015,863
faults min/maj0 / 0396 / 0
peak RSS5.2 GiB5.2 GiB
disk written00
disk read00
recall unequal: 1.0000 vs 0.9893; §7.4 compares at equal recall filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set

W12-sel1-ef32 filtered recall control, one keyword, ef=32 (latency only) parity

ef=32 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second6,4196,432parity
recall@101.00000.9964~equal
client p50293 µs304 µs1.04x
client p99444 µs399 µs0.90x
client p99.9493 µs455 µs0.92x
server p50205 µs203 µs~equal
server p99333 µs278 µs0.84x
wall clock7.8 s7.8 s~equal
queries sent50,00050,000
cpu12.8 s15.5 s1.22x
cpu, % of wall163%198%
waiting for a core2.6 ms10.7 ms4.05x
migrations09,702
faults min/maj0 / 061 / 0
peak RSS5.2 GiB5.2 GiB~equal
disk written00
disk read00
within the ±6.8% band this dataset's noise floor puts on W12-sel1-ef32: no measured difference, not a small one

W12-sel1-ef64 filtered recall control, one keyword, ef=64 (latency only) 1.19x

ef=64 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second6,2695,2581.19x
recall@101.00000.9997~equal
client p50302 µs376 µs1.24x
client p99448 µs448 µs~equal
client p99.9493 µs531 µs1.08x
server p50215 µs277 µs1.29x
server p99336 µs331 µs~equal
wall clock8.0 s9.5 s1.19x
queries sent50,00050,000
cpu13.0 s19.1 s1.46x
cpu, % of wall162%200%
waiting for a core2.0 ms12.6 ms6.20x
migrations09,517
faults min/maj0 / 065 / 0
peak RSS5.2 GiB5.2 GiB~equal
disk written00
disk read00

W12-sel1-ef128 filtered recall control, one keyword, ef=128 (latency only) 1.55x

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second5,9963,8621.55x
recall@101.00001.0000~equal
client p50310 µs513 µs1.66x
client p99458 µs593 µs1.30x
client p99.9495 µs679 µs1.37x
server p50222 µs413 µs1.86x
server p99344 µs476 µs1.39x
wall clock8.4 s13.0 s1.55x
queries sent50,00050,000
cpu13.5 s26.0 s1.92x
cpu, % of wall162%200%
waiting for a core2.0 ms17.1 ms8.67x
migrations09,573
faults min/maj0 / 064 / 0
peak RSS5.2 GiB5.2 GiB~equal
disk written00
disk read00

W12-sel1-ef256 filtered recall control, one keyword, ef=256 (latency only) 2.33x

ef=256 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second6,2712,6872.33x
recall@101.00001.0000~equal
client p50302 µs740 µs2.45x
client p99452 µs830 µs1.84x
client p99.9498 µs936 µs1.88x
server p50214 µs638 µs2.98x
server p99339 µs712 µs2.10x
wall clock8.0 s18.6 s2.33x
queries sent50,00050,000
cpu13.0 s37.5 s2.89x
cpu, % of wall162%201%
waiting for a core2.5 ms21.4 ms8.54x
migrations09,041
faults min/maj0 / 0199 / 0
peak RSS5.2 GiB5.2 GiB~equal
disk written00
disk read00

W12-sel1-ef512 filtered recall control, one keyword, ef=512 (latency only) 3.66x

ef=512 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second6,2281,7013.66x
recall@101.00001.0000~equal
client p50304 µs1.17 ms3.85x
client p99452 µs1.32 ms2.91x
client p99.9497 µs1.40 ms2.82x
server p50216 µs1.07 ms4.93x
server p99339 µs1.20 ms3.54x
wall clock8.1 s29.4 s3.65x
queries sent50,00050,000
cpu13.1 s59.4 s4.53x
cpu, % of wall162%202%
waiting for a core2.4 ms30.5 ms12.48x
migrations013,852
faults min/maj0 / 0435 / 0
peak RSS5.2 GiB5.2 GiB~equal
disk written00
disk read00

W12-sel10-ef32 filtered recall control, any of 10 keywords, ef=32 (latency only)

ef=32 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second3,1724,685-
recall@100.97770.8007
client p50625 µs422 µs
client p99861 µs534 µs
client p99.9958 µs603 µs
server p50532 µs319 µs
server p99750 µs423 µs
wall clock15.8 s10.7 s
queries sent50,00050,000
cpu28.1 s21.7 s
cpu, % of wall178%202%
waiting for a core1.7 ms9.0 ms
migrations011,149
faults min/maj0 / 0166 / 0
peak RSS5.2 GiB5.2 GiB
disk written00
disk read00
recall unequal: 0.9777 vs 0.8007; §7.4 compares at equal recall

W12-sel10-ef64 filtered recall control, any of 10 keywords, ef=64 (latency only)

ef=64 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,0983,377-
recall@100.99540.9331
client p50972 µs591 µs
client p991.10 ms717 µs
client p99.91.39 ms790 µs
server p50880 µs488 µs
server p991.00 ms605 µs
wall clock23.9 s14.8 s
queries sent50,00050,000
cpu44.2 s30.1 s
cpu, % of wall185%203%
waiting for a core1.2 ms10.2 ms
migrations013,038
faults min/maj0 / 0135 / 0
peak RSS5.2 GiB5.2 GiB
disk written00
disk read00
recall unequal: 0.9954 vs 0.9331; §7.4 compares at equal recall

W12-sel10-ef128 filtered recall control, any of 10 keywords, ef=128 (latency only)

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,3582,226-
recall@101.00000.9893
client p501.49 ms902 µs
client p991.61 ms1.07 ms
client p99.91.72 ms1.14 ms
server p501.39 ms796 µs
server p991.50 ms961 µs
wall clock36.9 s22.5 s
queries sent50,00050,000
cpu69.8 s45.5 s
cpu, % of wall189%202%
waiting for a core2.6 ms10.5 ms
migrations015,511
faults min/maj0 / 0250 / 0
peak RSS5.2 GiB5.2 GiB
disk written00
disk read00
recall unequal: 1.0000 vs 0.9893; §7.4 compares at equal recall

W12-sel10-ef256 filtered recall control, any of 10 keywords, ef=256 (latency only) 1.02x

ef=256 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,3581,325~equal
recall@101.00000.9992~equal
client p501.49 ms1.52 ms~equal
client p991.61 ms1.82 ms1.13x
client p99.91.75 ms1.93 ms1.10x
server p501.39 ms1.40 ms~equal
server p991.50 ms1.69 ms1.13x
wall clock36.9 s37.8 s1.02x
queries sent50,00050,000
cpu69.8 s76.1 s1.09x
cpu, % of wall190%202%
waiting for a core5.0 ms31.4 ms6.33x
migrations018,323
faults min/maj0 / 0407 / 0
peak RSS5.2 GiB5.2 GiB~equal
disk written00
disk read00

W12-sel10-ef512 filtered recall control, any of 10 keywords, ef=512 (latency only) 1.70x

ef=512 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,3597981.70x
recall@101.00001.0000~equal
client p501.49 ms2.52 ms1.69x
client p991.61 ms3.00 ms1.86x
client p99.91.71 ms3.16 ms1.85x
server p501.39 ms2.39 ms1.72x
server p991.50 ms2.87 ms1.91x
wall clock36.8 s62.7 s1.70x
queries sent50,00050,000
cpu69.8 s126.1 s1.81x
cpu, % of wall190%201%
waiting for a core4.8 ms57.6 ms12.04x
migrations018,614
faults min/maj0 / 01,068 / 0
peak RSS5.2 GiB5.2 GiB~equal
disk written00
disk read00

W13 scroll / pagination 1.10x

closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=200,000
metricstrawmANNQdrant
queries/second51,19246,5601.10x
client p50137 µs158 µs1.15x
client p99230 µs292 µs1.27x
client p99.9270 µs441 µs1.63x
server p5011 µs28 µs2.56x
server p9916 µs90 µs5.58x
wall clock4.0 s2.2 s0.55x
queries sent200,000100,000
cpu5.6 s9.1 s1.64x
cpu, % of wall138%411%
waiting for a core3.2 ms864.2 ms269.26x
migrations0112,331
faults min/maj0 / 0708 / 0
peak RSS5.2 GiB5.2 GiB~equal
disk written00
disk read00
strawmANN: -n 200000 (4x the table's 50000) Qdrant: -n 100000 (2x the table's 50000) strawmANN: 0.13 ms of the client's 0.14 ms p50 is not server time (12x), so 92% of this row is the load generator and the socket Qdrant: 0.13 ms of the client's 0.16 ms p50 is not server time (6x), so 82% of this row is the load generator and the socket 1.10x is 3.27x less work per query x 0.33x cores busy during the row (1.39 against 4.22) x 1.00x clock x 1.02x counted on-CPU share: the middle term is occupancy, not search speed

W11-steady mixed read/write below the rebuild threshold: search bench2 while 50,000 synthetic points append

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second18,2285,537-
client p50394 µs1.26 ms
client p991.96 ms4.79 ms
client p99.93.01 ms9.92 ms
server p50316 µs963 µs
server p991.11 ms4.08 ms
wall clock2.9 s9.1 s
queries sent50,00050,000
cpu18.9 s88.5 s
cpu, % of wall662%968%
waiting for a core1,345.8 ms54,590.2 ms
migrations0116,172
faults min/maj229,802 / 0936,740 / 224,007
peak RSS5.2 GiB8.7 GiB
disk written24.5 MiB3.5 GiB
disk read0365.6 MiB
strawmANN: qps over the 2.7 s the append ran (18,221 over the whole search); write overlap 11%; search covered 11% of the append; append 2,000 points/s Qdrant: qps over the 9.0 s the append ran (5,533 over the whole search); write overlap 36%; search covered 37% of the append; append 2,000 points/s search-during-write; no recall join; the search saw under 90% of either append, so its qps is the append's start and not all of it

W11 mixed read/write: search bench2 while 200,000 synthetic points append (runs last)

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second8,0332,387-
client p50430 µs2.62 ms
client p994.09 ms14.09 ms
client p99.96.06 ms25.83 ms
server p50348 µs2.15 ms
server p993.36 ms13.04 ms
wall clock6.3 s21.1 s
queries sent50,00050,000
cpu49.5 s267.7 s
cpu, % of wall781%1,271%
waiting for a core17,678.2 ms135,269.8 ms
migrations111108,738
faults min/maj810,544 / 01,842,780 / 391,257
peak RSS5.2 GiB14.7 GiB
disk written98.1 MiB6.6 GiB
disk read0736.2 MiB
strawmANN: qps over the 6.2 s the append ran (8,029 over the whole search); write overlap 10%; search covered 10% of the append; append 3,300 points/s Qdrant: qps over the 20.9 s the append ran (2,386 over the whole search); write overlap 35%; search covered 35% of the append; append 3,300 points/s search-during-write; no recall join; the search saw under 90% of either append, so its qps is the append's start and not all of it
Host discipline the §7.1 gate, per check, and ambient load per row

Ambient load per row

The §7.1 gate checks the machine once, at the start. Colour is the engine, as everywhere else; a hatched bar with a red edge is a row that had another process on the box while it ran, which the gate cannot see.

AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical) gate pass

One machine: both runs were measured on it, and the checks below are the same for both. What differed at each run start is listed underneath, per run.
okAMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical)
okgovernor=performance
okboost disabled
oksmt=on
declared rather than required (§7.1 asks for SMT on or off, not
for off), and hashed, so these rows never share a chart with
SMT-off ones
cpu N and cpu N+12 share a core (12 pairs);
a --server-cpus or --client-cpus set holding both cpus of a
pair measures the engine on half the cores it names
okprofile=as-deployed (isolcpus=none nohz_full=none)
the scheduler is running as it ships, so these numbers describe a
deployment rather than the engine in isolation; they may not be
compared against an `isolated` run, and the environment hash
enforces that
okthp=madvise numa_nodes=1 numa_balancing=0
okperf_event_paranoid=-1
oktask_delayacct=1, so time blocked on a device is measurable
okperf at /usr/bin/perf, so `workloads.py run --perf` can attach
okcgroup io controller reaches the engines' own scope, so block-layer read/write operations are measurable
okzig=0.16.0
busy: Xorg(5%)
measurement profile: as-deployed
strawmANN at run start · environment hash 373e8531fd3d1bf6
Qdrant at run start · environment hash 373e8531fd3d1bf6
What this does not establish

No Qdrant developer has used this tool. Nothing here has been run, checked against something already known, or disagreed with by anyone outside the project: every number has one author and one reviewer, and they are the same person. A result you can refute is more useful to us than one you accept.

It has never been pointed at a real Qdrant regression. The harness detects a known ISA slowdown and correctly reports no difference on a control row. That is internal consistency, not evidence it would flag a regression in your tree or stay quiet through a refactor. bench/harness/qdrant_ab.py exists to run two of your commits through it blind, and that experiment has not been done.

Part of the measured throughput gap is kernel width, not architecture. Read from a dev checkout: Qdrant's distance path tops out at four 256-bit accumulators and has no AVX-512, while these kernels use 512-bit ones on a machine that has them. That is a real difference and it is not the same claim as "the design is faster".

The engines are not equivalent, by construction. §1 removes sharding, replication, consensus, snapshots, sparse vectors, multivectors and disk-resident operation — permanently. Payload storage and filtering were in that list until 2026-09-03 and are now built (M7), so W12 measures a filtered search with a keyword index on both engines. Anything still unbuilt answers UNIMPLEMENTED naming the construct and reports n/a, never a silent degradation. A strawman that was not faster would mean it was badly built; the question is by how much, and where the model was wrong.

Reproducing this

Everything below runs from a clone. The engine has no dependencies; the analysis path is a uv project pinned by bench/uv.lock.

scripts/doctor.py                    # what can this machine measure?
bench/setup.py check                 # §7.1 host gate: governor, boost, isolation, idle
conformance/datasets/datasets.py fetch

# one engine at a time, §7.1
zig build -Doptimize=ReleaseFast
# --connections matters: W4 opens ~32 sockets (-t 16 -c 2) and a server with
# fewer closes the excess before the HTTP/2 preface, which the client reports
# only as "transport error".
./zig-out/bin/strawmann --port 6334 --connections 64
bench/harness/workloads.py run http://localhost:6334 strawmann --sink --report

# the other half of every number: recall, and the licence to compare at all
cd conformance
cargo run --release -- relevance --engine http://localhost:6334 --label strawmann \
  --base $DATA/sift1m.fbin --queries $DATA/sift1m_query.fbin \
  --ground-truth $DATA/gt/sift1m.euclid.k100.gt.json --metric euclid \
  --ef 32 64 128 256 512 --json ../bench/results/strawmann/recall.json
cargo run --release -- differ --strawmann http://localhost:6334 --qdrant http://localhost:6434 \
  --base $DATA/sift1m.fbin --queries $DATA/sift1m_query.fbin \
  --ground-truth $DATA/gt/sift1m.euclid.k100.gt.json \
  --json ../bench/results/strawmann/conformance.json

uv run --project bench bench/harness/report.py strawmann qdrant

Disagreements are the point. docs/spec.md is what all of this is measured against, docs/bugs.md lists the measurement bugs found so far (each produced a plausible wrong number rather than a failure), and docs/validation.md records what has and has not been checked.