strawmANN benchmark report

strawmANN against Qdrant · AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical) · generated 2026-10-03 11:42:56 CEST

Dataset
laion-small-clip
Vectors
100,000 × 512 · 5,000 held-out queries
Metric
cosine
Storage datatype
strawmANN float32 · Qdrant float32 (default)
Quantization
none (fp32) · separate rows measure binary (1 bit/dim), PQ (product), SQ8 scalar, turboquant
strawmANN: 55/55 ok Qdrant: 55/55 ok

Summary

At equal recall, strawmANN serves 1.53x to 1.81x Qdrant's throughput. That range spans the recall levels measured; allowing for the uncertainty in each recall figure widens it to 0.86x to 3.03x.

throughput at equal recall
1.53 to 1.81x
strawmANN over Qdrant, recall@10 0.963 to 0.998
p99 latency at 90% load
strawmANN2.14 ms
Qdrant4.91 ms
W4-sat90, open loop, one offered rate
upload and index build
strawmANN9 s
Qdrant14 s
W1 + W2
peak memory (RSS)
strawmANN1.5 GiB
Qdrant3.1 GiB
largest before the concurrent-write rows (W11); of it, anonymous 315.8 MiB / 288.5 MiB
written to disk
strawmANN2.3 GiB
Qdrant10.1 GiB
before the concurrent-write rows (W11); they wrote 49.3 MiB / 9.5 GiB more

Throughput against recall

Read it vertically: at any recall both engines reach, the higher curve is faster. Throughput from bfb W10, recall from the conformance sweep, joined on ef. ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.

Every figure is the median of 3 passes per engine, run alternately (A/B/A/B). What this does not establish.

Glossary
recall@10
Of the ten nearest neighbours a query really has, the share the engine returned. 1.0 is a perfect answer. "Really has" is settled by an exhaustive fp64 search, not by the other engine.
ef
The candidate-list size: how many nodes the index keeps in play while searching. Larger is slower and more accurate. It does not mean the same amount of work in both engines, which is why the headline holds recall fixed instead.
MRDE
Mean relative distance error: when the engine returns a wrong neighbour, how much further away it is than the right one. Small numbers mean the misses were near-misses.
qps
Queries per second. On a batched row one request carries several queries, and the request rate is shown under it.
p50 / p99 / p99.9
The latency half of requests beat, that 99% beat, that 99.9% beat. The tail is what a user notices.
closed / open loop
A closed loop sends the next request only when the last one comes back, so a slow server receives less work and its tail looks better than it is. An open loop sends at a fixed rate regardless.
noise floor
How much a number moves between identical runs on this machine. A ratio inside it is shown grey with ≈: no measured difference. Hover a ratio to see its band.
conformance tier
What the two engines were shown to agree on before any speed was quoted. T1 licenses a single engine's own numbers; T3 licenses comparing the two, and requires their recall to be statistically indistinguishable.
§ numbers
Sections of docs/spec.md, the written rule each claim in the appendix is measured against. §7.1 is the host gate, §7.4 the comparison rules, §8 what may be published.

Throughput

Queries per second; higher is better. The ratio is strawmANN over Qdrant: green is faster, red slower, grey ≈ inside the noise floor (hover a ratio for its band). A dash means the pair is not compared, and the note says why. An ef sweep is one row showing its range; every measurement is under All throughput rows.

workloadstrawmANNQdrantrationotes
W3search, fp32, single query1,6461,637parity
W4search, saturating (closed loop)10,80010,924parity
W5search batched (16 distinct dataset queries per request)10,643664 requests/s3,754234 requests/s2.83x
W6quantized: scalar5,5074,4431.24x
W6 ef 32 to 512SQ8 recall control (latency only)2,072 to 10,8471,858 to 7,5021.13 to 1.41x2 of 5 points not compared recall differs one pass apart: Qdrant pass 3 -21%
W7quantized: binary + oversampling5,7694,6161.25x
W7-2bitquantized: binary quantization, 2 bits + oversampling5,4654,6361.18x
W7-1p5bitquantized: binary quantization, 1.5 bits + oversampling5,5254,5381.22x
W8quantized: PQ3,4402,8881.19x
W14-1bitquantized: TurboQuant, 1 bit + oversampling4,6124,570parity
W14-1p5bitquantized: TurboQuant, 1.5 bits + oversampling4,3564,0901.07x
W14-2bitquantized: TurboQuant, 2 bits + oversampling5,4644,5961.19x
W14-4bitquantized: TurboQuant, 4 bits + oversampling5,0864,5961.11x
W9exact / brute force311311parity
W10 ef 32 to 512recall control (latency only)3,423 to 27,9013,373 to 18,0651.12 to 1.42x1 of 5 points at parity 1 of 5 points not compared recall differs
W12-sel1filtered search, one keyword (~1% of bench12)7,0963,368-one pass apart: strawmANN pass 1 -16%
W12-sel10filtered search, any of 10 keywords (~10% of bench12)1,3221,803-recall differs: 1.0000 vs 0.9854
W12-sel1 ef 32 to 512filtered recall control, one keyword (latency only)6,833 to 7,3121,693 to 5,7741.26 to 4.04x1 of 5 points not compared one pass apart: strawmANN pass 1 -39%
W12-sel10 ef 32 to 512filtered recall control, any of 10 keywords (latency only)1,319 to 3,212612 to 3,9801.29 to 2.16x3 of 5 points not compared recall differs
W13scroll / pagination50,28246,2731.09xstrawmANN: 92% is client and socket, not server Qdrant: 82% is client and socket, not server
W11-steadymixed read/write below the rebuild threshold: search bench2 while 5,000 synthetic points append10,1027,170-search-during-write; no recall join
W11mixed read/write: search bench2 while 20,000 synthetic points append (runs last)9,5267,151-search-during-write; no recall join

Throughput by workload

Higher is better. Hatched bars are rows the table does not compare (W11): each bar is that engine's own rate, and the pair is not a result.

Recall

recall@10 against an exact fp64 search, not against the other engine. Equal ef is not equal work in the two engines, so the comparison that counts holds recall fixed and compares throughput there.

At matched recall

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.9630strawmANN at ef=32, Qdrant interpolated27,90115,3911.81x1.66 – 1.98x
0.9786Qdrant at ef=64, strawmANN interpolated20,24212,4381.63x1.48 – 1.79x
0.9852strawmANN at ef=64, Qdrant interpolated17,6679,9251.78x1.57 – 2.01x
0.9906Qdrant at ef=128, strawmANN interpolated12,8408,2681.55x1.28 – 1.88x
0.9940strawmANN at ef=128, Qdrant interpolated10,4876,1661.70x1.39 – 2.09x
0.9956Qdrant at ef=256, strawmANN interpolated8,2375,3801.53x1.12 – 2.10x
0.9976strawmANN at ef=256, Qdrant interpolated6,0123,5281.70x1.25 – 2.32x
0.9979Qdrant at ef=512, strawmANN interpolated5,4313,3731.61x0.86 – 3.03x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W10 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column is why the ratio is not a result on its own. The ratio treats the anchor's recall as exact; it is an estimate, and moving it across its 95% interval moves the interpolated rate with it — near the top of the sweep 0.003 of recall spans a factor of 1.8, so two decimals there quote the interpolation. Even that is the narrow reading: it moves one recall and not the other, adjacent anchors share bracketing segments, and the throughputs behind it are medians of the run's passes.

At matched recall, filtered to 10%

The same reading under a keyword filter over the points the condition matched, 9,943 to 10,109 across 3 builds in strawmANN and 9,917 to 10,066 across 3 builds in Qdrant (bfb draws the keyword payloads unseeded at each upload). The per-row table refuses W12-sel10 a ratio because the two engines land just outside the recall band at the one ef it measures; held at equal recall instead, the comparison exists at every recall both engines reach. Note the shape rather than any single number: one engine's curve is flat in ef and the other's is steep, so where you match decides the ratio.

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.9808strawmANN at ef=32, Qdrant interpolated3,2121,8691.72x1.65 – 1.78x
0.9854Qdrant at ef=128, strawmANN interpolated2,7361,8001.52x1.33 – 1.73x
0.9932strawmANN at ef=64, Qdrant interpolated2,0771,2391.68x1.49 – 1.89x
0.9972Qdrant at ef=256, strawmANN interpolated1,5881,0241.55x1.39 – 1.73x
0.9993Qdrant at ef=512, strawmANN interpolated1,3796122.25x2.13 – 2.38x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W12-sel10 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column reads as it does in the fp32 table above.

Recall by ef

efstrawmANNQdrant
recall@1recall@10recall@100MRDErecall@1recall@10recall@100MRDE
ef=320.97240.9630-1.12e-030.96760.9513-1.40e-03
ef=640.98720.9852-3.99e-040.98020.9786-6.50e-04
ef=1280.99420.99400.97551.51e-040.98940.99060.96772.65e-04
ef=2560.99780.99760.99126.03e-050.99460.99560.98811.27e-04
ef=5120.99900.99890.99662.38e-050.99760.99790.99536.26e-05
How recall was measured

5,000 held-out queries, limit 10, ε=9.537e-07, against the fp64 oracle rather than against the other engine. Base checksum 7163721f6ee7759a: the same corpus the latency rows were measured on. Each pass rebuilds the collection. Across 3 independent builds of it, recall@10 at ef=512 spread by: strawmANN 0.00008; Qdrant 0.00006. Over the same builds the engine reported unreachable nodes: strawmANN 0 of 100,000, in-degree zero 4 — a node nothing points at is invisible at any ef, so that is where a graph-quality difference shows rather than being inferred from the recall beside it. Level seed: strawmANN 0x57ea3111. Recorded as provenance: under this engine's level draw, six builds at four seeds sit within 0.00003 of recall@10 at ef 512, well inside the spread above.

What each quantization costs in recall

Colour is the encoding here, not the engine — the engines are the line style. Each sweep is joined on the quantization parameters it actually sent, so a curve speaks only for the search its row ran. Higher is better, and the encodings are not free: read this against the throughput their rows bought.

Latency

Client-side round trip. A closed loop understates the tail (a stalled server stops receiving requests), so this keeps the fixed-rate, open-loop rows, which offer the same load to both engines, plus single-client W3 and batched W5. Every row is under All latency rows.

workloadstrawmANNQdrant
p50p99p99.9p50p99p99.9
W3 search, fp32, single query574 µs966 µs1.07 ms611 µs790 µs860 µs
W4-sat50 search, fixed rate at 50% of saturation (open loop)696 µs1.34 ms1.59 ms856 µs1.95 ms2.39 ms
W4-sat70 search, fixed rate at 70% of saturation (open loop)861 µs1.89 ms2.21 ms942 µs2.20 ms3.30 ms
W4-sat90 search, fixed rate at 90% of saturation (open loop)913 µs2.14 ms4.15 ms1.33 ms4.91 ms8.14 ms
W5 search batched (16 distinct dataset queries per request)3.00 ms3.52 ms3.69 ms8.50 ms9.40 ms9.87 ms

Ingest and index build

Qdrant indexes while it ingests, so its upload time already contains most of the indexing; strawmANN uploads raw and builds afterwards. Compare the sum, not the upload line.

Where the load time goes

Solid is upload, hatched is the wait for Green. The bar's whole length is the sum this section asks you to compare; the split is why the upload line alone is not comparable between these engines.

workloadstrawmANNQdrant
W0-uploadd=4 floor: graph traversal with the distance taken out4.01 s3.01 s †
W1ingest throughput (no index wait)0.29 s1.11 s
W2index build, time to Green9.03 s13.03 s
W6-uploadsearch, SQ8 scalar quantization8.03 s8.03 s
W7-uploadsearch, binary quantization5.02 s8.03 s
W7-2bit-uploadsearch, binary quantization6.02 s8.03 s
W7-1p5bit-uploadsearch, binary quantization6.02 s9.03 s
W8-uploadsearch, product quantization45.15 s45.12 s
W14-1bit-uploadupload10.03 s9.03 s
W14-1p5bit-uploadupload10.03 s9.03 s
W14-2bit-uploadupload10.03 s9.03 s
W14-4bit-uploadupload10.03 s9.03 s
W12-uploadupload10.03 s21.06 s
W1 + W2upload and index, together9.3 s14.1 s1.52x

† at the load generator's polling floor. bfb decides a collection is Green by polling once a second and requiring three consecutive Green replies, having slept a second before the first poll, so no Time-to-Green it can report is below 3 s regardless of how fast the build was. A marked cell is an upper bound on the build and a measurement of the polling loop. The bias is a constant added to both engines, so the difference between them survives it and the ratio does not — and it is the faster engine the ratio understates.

Memory and disk

What each engine held in memory and moved to and from disk over the whole run. Per-row figures are under Storage and I/O.

strawmANNQdrant
storage on disk1.2 GiBallocated at --capacity1.3 GiB
peak memory (RSS)1.5 GiB3.1 GiB
read from disk0416.0 MiB
written to disk2.3 GiB10.1 GiB

Peak memory and storage on disk are read before the concurrent-write rows (W11): an engine rewriting segments maps old and new files at once, and RSS counts each mapping. Bytes written are summed over the same rows, since W11's volume is set by the harness's write rate; the other disk figures are totals over every row.

Appendix

How the run was set up, every row of every table, and the diagnostics behind them. Charts on a logarithmic axis say so on the axis: read the positions there, not the distances.

Run conditions and conformance

Measured as-deployed: the scheduler was left as it ships, so a difference is what a user would see rather than the engine in isolation, and ambient load is part of the measurement. Qdrant ran equal-work: asked for one populated graph (--segments 1), as strawmANN serves, so ef means the same thing on both sides. This is the configuration §8's comparative licensing is built around, and it is a control rather than a deployment. Its optimizer treats the count as a target; what it actually held is read back from the engine and stated under Collection config.

Conformance tier reached: T4 quantization fidelity. T1 passed, so single-engine performance rows are licensed (§8). T3 passed, so a strawmANN-vs-Qdrant throughput comparison is licensed: the engines are at equal recall (§7.4). Conformance hash c7c43624d600016f.

On exact search the two engines' scores differ by at most 2.980e-07 (p99 1.788e-07), against a calibrated absolute ε of 9.537e-07. §8.1: bit-exactness is unachievable between two different summation orders, so this — equal values within a measured tolerance — is the claim the speed rests on.

The conformance harness's single-client rate agrees on the ordering (2.37x to 2.72x); a different instrument, so no ratio is formed from it.

Every figure is the median of 3 passes per row per engine, alternated A/B/A/B (§7.4), and the spread is this run's own over 42 of 55 rows. The ingest and index-build rows (W0-upload, W1, W2, W6-upload, W7-upload, W7-2bit-upload, W7-1p5bit-upload, W8-upload, W14-1bit-upload, W14-1p5bit-upload, W14-2bit-upload, W14-4bit-upload, W12-upload) have none, so no build time carries a verdict. Each pass rebuilds the index, so the bands include build variance and are wider than a floor measured against one standing graph.

Units are queries per second. bfb reports rps, which counts batch requests: at --search-batch-size 16 the two differ by 16x, and reading one as the other once turned a 2.2x speedup into an apparent 7x regression. The tables show the request rate wherever it diverges.

The open-loop rows are not a speed. W4-sat50/70/90 offer a fixed fraction of measured saturation; serving it means the engine kept up, not that it was faster, so no ratio is printed. They exist because --parallel is a closed loop, where a stalled server stops receiving requests and understates its own tail.

ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.

What was measured builds, dataset, host
strawmANN
sm-laion-perf-1003
Qdrant
qd-laion-perf-1003
enginecommit f346aea94dbf
ReleaseFast, native build, vnni on
binary sha256 aca46d195c7e22bc
built 2026-10-02T14:42:33Z with zig 0.16.0
version 1.19.2-dev
binary ~/.cache/strawmann/qdrant-dbeb0f73/qdrant
sha256 dbeb0f73dea2d371 build 878843e6 (from the server's banner, not a checkout)
native binary outside a checkout: sha256 is the identity (§8.9)
network native no container in the path, like strawmann
measured2026-10-03T06:30:49Z to 2026-10-03T08:25:46Z (3 passes)2026-10-03T06:54:01Z to 2026-10-03T08:50:18Z (3 passes)
profileas-deployed · both engines
§7.1 gatepass · both engines
environment hash373e8531fd3d1bf6 · both engines
load generatorbfb dev @ 29240511 (qdrant/bfb#183; carries #172, findings 32's --rps reaping fix) · both engines
clientqdrant-client 1.16.1-dev (git dev branch) · both engines
launched as~/Workspace/strawmann/zig-out/bin/strawmann --port 6334 --capacity 125000 --connections 64 --workers 7 --io-threads 1 --pin --cpus 4-11 --data-dir ~/.cache/strawmann/strawmann-storage --default-placement cached~/.cache/strawmann/qdrant-dbeb0f73/qdrant
cores the engine could use4-11 (8 cores) requested by the harness
4-11 (8 cores) observed on 8 threads
0-23 (24 cores) observed on 1 thread
4-11 (8 cores) requested by the harness
4-11 (8 cores) observed on 52 threads

Dataset

laion-small-clip, 100000 × 512, cosine, 5000 held-out queries
ground truth: shipped, and diffed against our fp64 recompute
1 file, each pinned by sha256 in datasets.json

Host

AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M
24 logical cores · 58.6 GB · kernel 7.0.0-38-generic
memory bandwidth: 75.7 GB/s aggregate (24 threads) · 45.2 GB/s single core (60% of bus)
strawmann startup banner, 2026-10-03T06:30:49Z; the Qdrant run agrees within 10%
All throughput rows every measurement, with its notes
workloadstrawmANNQdrantrationotes
W0d=4 floor: graph traversal with the distance taken out4,8244,0751.18x1.18x is 1.60x less work per query x 0.66x cores busy during the row (0.66 against 1.00) x 1.00x clock x 1.12x counted on-CPU share: the middle term is occupancy, not search speed d=4 floor: this measures graph traversal and heaps with the distance taken out, not request plumbing (findings 28: 76% of the server's CPU was HNSW search, 24% kernel; the RPC path is a small remainder)
W3search, fp32, single query1,6461,637paritywithin the ±2.3% band this dataset's noise floor puts on W3: no measured difference, not a small one
W4search, saturating (closed loop)10,80010,924paritywithin the ±2.0% band this dataset's noise floor puts on W4: no measured difference, not a small one
W4-sat50search, fixed rate at 50% of saturation (open loop)5,4895,489offered
W4-sat70search, fixed rate at 70% of saturation (open loop)7,6857,685offered
W4-sat90search, fixed rate at 90% of saturation (open loop)9,8799,879offered
W5search batched (16 distinct dataset queries per request)10,643664 requests/s3,754234 requests/s2.83xper-batch latency (16 queries/request); queries: dataset, random-sample 2.83x is 0.78x less work per query x 3.58x cores busy during the row (7.04 against 1.97) x 1.00x clock x 1.01x counted on-CPU share: the middle term is occupancy, not search speed batched 16, so qps and rps differ by 16x; the ratio uses qps
W6quantized: scalar5,5074,4431.24x
W6-ef32SQ8 recall control, ef=32 (latency only)10,8477,502-recall unequal: 0.9624 vs 0.9504; §7.4 compares at equal recall; the rescore pools match, so the encoders differ
W6-ef64SQ8 recall control, ef=64 (latency only)8,4525,9881.41x1.41x is 1.72x less work per query x 0.77x cores busy during the row (1.51 against 1.97) x 1.00x clock x 1.07x counted on-CPU share: the middle term is occupancy, not search speed
W6-ef128SQ8 recall control, ef=128 (latency only)5,4024,4691.21x
W6-ef256SQ8 recall control, ef=256 (latency only)3,4103,0251.13x
W6-ef512SQ8 recall control, ef=512 (latency only)2,0721,858-Qdrant pass 3 -21% against two passes that agree: one pass set the spread on this row, so a band built from it is not a noise band
W7quantized: binary + oversampling5,7694,6161.25x
W7-2bitquantized: binary quantization, 2 bits + oversampling5,4654,6361.18x
W7-1p5bitquantized: binary quantization, 1.5 bits + oversampling5,5254,5381.22x
W8quantized: PQ3,4402,8881.19x
W14-1bitquantized: TurboQuant, 1 bit + oversampling4,6124,570paritywithin the ±2.0% band this dataset's noise floor puts on W14-1bit: no measured difference, not a small one
W14-1p5bitquantized: TurboQuant, 1.5 bits + oversampling4,3564,0901.07x
W14-2bitquantized: TurboQuant, 2 bits + oversampling5,4644,5961.19x
W14-4bitquantized: TurboQuant, 4 bits + oversampling5,0864,5961.11x
W9exact / brute force311311paritywithin the ±12.2% band this dataset's noise floor puts on W9: no measured difference, not a small one brute force over the whole collection, no index involved
W10-ef32recall control, ef=32 (latency only)27,90118,065-strawmANN: -n 100000 (2x the table's 50000) recall unequal: 0.9630 vs 0.9513; §7.4 compares at equal recall
W10-ef64recall control, ef=64 (latency only)17,66712,4381.42x
W10-ef128recall control, ef=128 (latency only)10,4878,2681.27x
W10-ef256recall control, ef=256 (latency only)6,0125,3801.12x
W10-ef512recall control, ef=512 (latency only)3,4233,373paritywithin the ±2.0% band this dataset's noise floor puts on W10-ef512: no measured difference, not a small one
W12-sel1filtered search, one keyword (~1% of bench12)7,0963,368-strawmANN pass 1 -16% against two passes that agree: one pass set the spread on this row, so a band built from it is not a noise band filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set
W12-sel10filtered search, any of 10 keywords (~10% of bench12)1,3221,803-recall unequal: 1.0000 vs 0.9854; §7.4 compares at equal recall filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set
W12-sel1-ef32filtered recall control, one keyword, ef=32 (latency only)7,2915,7741.26x
W12-sel1-ef64filtered recall control, one keyword, ef=64 (latency only)7,3124,5271.62x
W12-sel1-ef128filtered recall control, one keyword, ef=128 (latency only)7,2553,3392.17x2.17x is 2.72x less work per query x 0.80x cores busy during the row (1.60 against 2.00) x 1.00x clock: the middle term is occupancy, not search speed
W12-sel1-ef256filtered recall control, one keyword, ef=256 (latency only)7,0762,441-strawmANN pass 1 -39% against two passes that agree: one pass set the spread on this row, so a band built from it is not a noise band
W12-sel1-ef512filtered recall control, one keyword, ef=512 (latency only)6,8331,6934.04x4.04x is 5.32x less work per query x 0.79x cores busy during the row (1.59 against 2.00) x 1.00x clock x 0.96x counted on-CPU share: the middle term is occupancy, not search speed
W12-sel10-ef32filtered recall control, any of 10 keywords, ef=32 (latency only)3,2123,980-recall unequal: 0.9808 vs 0.8164; §7.4 compares at equal recall
W12-sel10-ef64filtered recall control, any of 10 keywords, ef=64 (latency only)2,0772,780-recall unequal: 0.9932 vs 0.9326; §7.4 compares at equal recall
W12-sel10-ef128filtered recall control, any of 10 keywords, ef=128 (latency only)1,3191,800-recall unequal: 1.0000 vs 0.9854; §7.4 compares at equal recall
W12-sel10-ef256filtered recall control, any of 10 keywords, ef=256 (latency only)1,3191,0241.29x
W12-sel10-ef512filtered recall control, any of 10 keywords, ef=512 (latency only)1,3206122.16x
W13scroll / pagination50,28246,2731.09xstrawmANN: -n 200000 (4x the table's 50000) Qdrant: -n 100000 (2x the table's 50000) strawmANN: 0.13 ms of the client's 0.14 ms p50 is not server time (13x), so 92% of this row is the load generator and the socket Qdrant: 0.13 ms of the client's 0.16 ms p50 is not server time (6x), so 82% of this row is the load generator and the socket 1.09x is 3.29x less work per query x 0.33x cores busy during the row (1.38 against 4.24) x 1.00x clock x 1.01x counted on-CPU share: the middle term is occupancy, not search speed
W11-steadymixed read/write below the rebuild threshold: search bench2 while 5,000 synthetic points append10,1027,170-strawmANN: qps over the 25.0 s of the append the search saw (10,152 over the whole search); append 200 points/s Qdrant: qps over the 25.0 s of the append the search saw (7,523 over the whole search); append 200 points/s search-during-write; no recall join
W11mixed read/write: search bench2 while 20,000 synthetic points append (runs last)9,5267,151-strawmANN: qps over the 66.7 s of the append the search saw (9,566 over the whole search); append 300 points/s Qdrant: qps over the 66.7 s of the append the search saw (7,528 over the whole search); append 300 points/s search-during-write; no recall join

W10: throughput against ef

ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.
All latency rows closed loop included, p50 to max

A row flagged “… not server time” is one where the client saw far more than the server reported: below the rate at which a queue can form, what is left is the load generator's own scheduling.

workloadstrawmANN p50strawmANN p95strawmANN p99strawmANN p99.9strawmANN maxQdrant p50Qdrant p95Qdrant p99Qdrant p99.9Qdrant max
W0d=4 floor: graph traversal with the distance taken out closed loop 205 µs232 µs250 µs305 µs3.61 ms242 µs274 µs300 µs364 µs2.19 ms
W3search, fp32, single query closed loop 574 µs888 µs966 µs1.07 ms2.69 ms611 µs747 µs790 µs860 µs3.62 ms
W4search, saturating (closed loop) closed loop 5.90 ms6.34 ms6.54 ms19.21 ms24.65 ms5.36 ms10.67 ms14.89 ms26.91 ms37.17 ms
W4-sat50search, fixed rate at 50% of saturation (open loop) open loop 696 µs1.10 ms1.34 ms1.59 ms4.44 ms856 µs1.56 ms1.95 ms2.39 ms6.06 ms
W4-sat70search, fixed rate at 70% of saturation (open loop) open loop 861 µs1.58 ms1.89 ms2.21 ms5.27 ms942 µs1.81 ms2.20 ms3.30 ms10.24 ms
W4-sat90search, fixed rate at 90% of saturation (open loop) open loop 913 µs1.68 ms2.14 ms4.15 ms6.03 ms1.33 ms3.02 ms4.91 ms8.14 ms19.53 ms
W5search batched (16 distinct dataset queries per request) closed loop 3.00 ms3.36 ms3.52 ms3.69 ms8.21 ms8.50 ms9.16 ms9.40 ms9.87 ms12.68 ms
W6quantized: scalar closed loop 355 µs463 µs538 µs660 µs1.36 ms447 µs502 µs534 µs632 µs1.71 ms
W6-ef32SQ8 recall control, ef=32 (latency only) closed loop 180 µs224 µs247 µs280 µs990 µs260 µs310 µs344 µs389 µs1.43 ms
W6-ef64SQ8 recall control, ef=64 (latency only) closed loop 229 µs296 µs324 µs369 µs975 µs327 µs389 µs437 µs491 µs1.45 ms
W6-ef128SQ8 recall control, ef=128 (latency only) closed loop 360 µs485 µs544 µs605 µs1.41 ms445 µs500 µs531 µs618 µs1.67 ms
W6-ef256SQ8 recall control, ef=256 (latency only) closed loop 576 µs766 µs868 µs950 µs1.41 ms662 µs745 µs779 µs842 µs1.96 ms
W6-ef512SQ8 recall control, ef=512 (latency only) closed loop 964 µs1.15 ms1.31 ms1.53 ms1.87 ms1.08 ms1.23 ms1.28 ms1.36 ms4.22 ms
W7quantized: binary + oversampling closed loop 341 µs425 µs469 µs629 µs1.51 ms423 µs514 µs550 µs607 µs1.76 ms
W7-2bitquantized: binary quantization, 2 bits + oversampling closed loop 361 µs441 µs474 µs537 µs1.26 ms420 µs514 µs548 µs609 µs1.65 ms
W7-1p5bitquantized: binary quantization, 1.5 bits + oversampling closed loop 357 µs441 µs477 µs541 µs1.06 ms431 µs526 µs563 µs621 µs1.81 ms
W8quantized: PQ closed loop 580 µs686 µs726 µs780 µs1.36 ms692 µs776 µs816 µs897 µs1.99 ms
W14-1bitquantized: TurboQuant, 1 bit + oversampling closed loop 433 µs507 µs538 µs604 µs1.17 ms429 µs520 µs554 µs610 µs1.62 ms
W14-1p5bitquantized: TurboQuant, 1.5 bits + oversampling closed loop 458 µs545 µs576 µs631 µs1.52 ms484 µs577 µs655 µs726 µs1.68 ms
W14-2bitquantized: TurboQuant, 2 bits + oversampling closed loop 362 µs444 µs476 µs519 µs1.13 ms425 µs518 µs551 µs607 µs3.75 ms
W14-4bitquantized: TurboQuant, 4 bits + oversampling closed loop 387 µs482 µs516 µs572 µs1.16 ms429 µs499 µs557 µs623 µs1.64 ms
W9exact / brute force closed loop 25.51 ms32.73 ms35.49 ms39.53 ms39.81 ms26.05 ms29.72 ms31.80 ms36.92 ms40.22 ms
W10-ef32recall control, ef=32 (latency only) closed loop 279 µs360 µs403 µs499 µs2.06 ms411 µs649 µs801 µs1.11 ms2.42 ms
W10-ef64recall control, ef=64 (latency only) closed loop 447 µs583 µs644 µs798 µs2.89 ms595 µs984 µs1.19 ms1.55 ms2.81 ms
W10-ef128recall control, ef=128 (latency only) closed loop 757 µs1.02 ms1.14 ms1.29 ms4.59 ms885 µs1.55 ms1.84 ms2.35 ms3.87 ms
W10-ef256recall control, ef=256 (latency only) closed loop 1.32 ms1.80 ms2.02 ms2.28 ms6.84 ms1.36 ms2.32 ms2.78 ms3.48 ms5.42 ms
W10-ef512recall control, ef=512 (latency only) closed loop 2.33 ms3.16 ms3.56 ms4.06 ms8.28 ms2.24 ms3.60 ms4.18 ms5.27 ms8.66 ms
W12-sel1filtered search, one keyword (~1% of bench12) closed loop 271 µs340 µs449 µs698 µs1.52 ms591 µs645 µs676 µs751 µs2.23 ms
W12-sel10filtered search, any of 10 keywords (~10% of bench12) closed loop 1.51 ms1.61 ms1.64 ms1.70 ms4.74 ms1.11 ms1.28 ms1.34 ms1.42 ms4.68 ms
W12-sel1-ef32filtered recall control, one keyword, ef=32 (latency only) closed loop 269 µs309 µs339 µs560 µs1.26 ms341 µs385 µs413 µs489 µs1.67 ms
W12-sel1-ef64filtered recall control, one keyword, ef=64 (latency only) closed loop 269 µs304 µs335 µs569 µs1.26 ms438 µs484 µs510 µs559 µs1.89 ms
W12-sel1-ef128filtered recall control, one keyword, ef=128 (latency only) closed loop 271 µs307 µs331 µs504 µs1.33 ms596 µs654 µs689 µs761 µs2.23 ms
W12-sel1-ef256filtered recall control, one keyword, ef=256 (latency only) closed loop 273 µs341 µs372 µs629 µs1.34 ms815 µs886 µs928 µs998 µs2.60 ms
W12-sel1-ef512filtered recall control, one keyword, ef=512 (latency only) closed loop 284 µs340 µs402 µs629 µs1.21 ms1.18 ms1.27 ms1.33 ms1.46 ms3.34 ms
W12-sel10-ef32filtered recall control, any of 10 keywords, ef=32 (latency only) closed loop 627 µs723 µs775 µs935 µs1.66 ms495 µs587 µs638 µs716 µs1.87 ms
W12-sel10-ef64filtered recall control, any of 10 keywords, ef=64 (latency only) closed loop 974 µs1.13 ms1.17 ms1.23 ms2.30 ms714 µs829 µs886 µs979 µs3.29 ms
W12-sel10-ef128filtered recall control, any of 10 keywords, ef=128 (latency only) closed loop 1.52 ms1.61 ms1.65 ms1.73 ms4.71 ms1.11 ms1.28 ms1.35 ms1.45 ms3.26 ms
W12-sel10-ef256filtered recall control, any of 10 keywords, ef=256 (latency only) closed loop 1.52 ms1.61 ms1.64 ms1.73 ms4.87 ms1.95 ms2.26 ms2.37 ms2.50 ms4.53 ms
W12-sel10-ef512filtered recall control, any of 10 keywords, ef=512 (latency only) closed loop 1.51 ms1.61 ms1.65 ms1.72 ms4.69 ms3.26 ms3.73 ms3.90 ms4.12 ms6.70 ms
W13scroll / pagination closed loop 139 µs207 µs235 µs275 µs1.05 ms159 µs224 µs296 µs440 µs2.28 ms
W11-steadymixed read/write below the rebuild threshold: search bench2 while 5,000 synthetic points append closed loop 771 µs1.04 ms1.34 ms3.00 ms10.06 ms932 µs1.71 ms2.26 ms3.83 ms24.44 ms
W11mixed read/write: search bench2 while 20,000 synthetic points append (runs last) closed loop 809 µs1.10 ms1.98 ms3.91 ms12.33 ms937 µs1.70 ms2.25 ms6.85 ms135.73 ms
Storage and I/O block layer and syscalls, per row

The syscall rows count every descriptor, sockets included, so on a search row they measure the network rather than the disk.

strawmANNQdrant
storage on disk1.2 GiBbefore W11-steady, W11allocated at --capacity1.3 GiBbefore W11-steady, W11
peak RSS1.5 GiB5.0 GiB3.1 GiB before the writers
of which anonymous315.8 MiB288.5 MiB
of which file-backed1.2 GiB2.2 GiB
disk read bytes0416.0 MiB
disk write bytes2.3 GiB19.6 GiB
disk read ops012,702
disk write ops16,702348,284
syscall reads (all fds)16,037,87113,742
syscall writes (all fds)7,446,2859,995,151

measured via strawmANN: proc+cgroup / Qdrant: proc+cgroup. proc supplies syscall counts and block-layer bytes, cgroup supplies block-layer operations and bytes, so a row one interface does not carry reads n/a via …. unknown means it was not measured, and 0 means it was: an engine started with no --data-dir has no store, which is the row this section exists for. storage on disk is the level before the rows with a concurrent writer: during those an engine that rewrites segments is caught mid-rewrite, and the same row has read 3.47 and 10.11 GiB on two runs of one binary. peak RSS does include them, being a peak.

Per workload

A search row doing block-layer reads is an engine going to disk to answer a query.

workloadstrawmANN readstrawmANN writtenstrawmANN read opsstrawmANN write opsQdrant readQdrant writtenQdrant read opsQdrant write ops
W0-upload06.1 MiB0010.2 MiB35.0 MiB375950
W0000990000
W10195.4 MiB00768.0 KiB436.5 MiB574,083
W20195.4 MiB0311.9 MiB843.1 MiB41114,911
W30003,1560006
W400000000
W4-sat5000000000
W4-sat7000000000
W4-sat9000000000
W500000000
W6-upload0195.4 MiB0012.7 MiB992.2 MiB43516,643
W60003,1530000
W6-ef3200000006
W6-ef6400030000
W6-ef12800000000
W6-ef25600000000
W6-ef51200000000
W7-upload0195.4 MiB0018.7 MiB973.9 MiB63017,414
W70003,1530000
W7-2bit-upload0195.4 MiB0014.1 MiB902.0 MiB47416,056
W7-2bit00030000
W7-1p5bit-upload0195.4 MiB0015.6 MiB903.3 MiB54316,239
W7-1p5bit00030000
W8-upload0195.4 MiB03,15411.6 MiB814.3 MiB35714,401
W800030006
W14-1bit-upload0195.4 MiB0016.1 MiB898.4 MiB52516,084
W14-1bit00000009
W14-1p5bit-upload0195.4 MiB0015.2 MiB880.6 MiB48315,638
W14-1p5bit00000000
W14-2bit-upload0195.4 MiB0012.8 MiB889.9 MiB42015,805
W14-2bit00000003
W14-4bit-upload0195.4 MiB0011.7 MiB912.4 MiB40216,137
W14-4bit00000006
W900000000
W10-ef3200000000
W10-ef6400000000
W10-ef12800000000
W10-ef25600000000
W10-ef51200000000
W12-upload0195.4 MiB0022.1 MiB837.4 MiB68415,314
W12-sel10003,1530000
W12-sel1000030000
W12-sel1-ef3200000000
W12-sel1-ef6400000000
W12-sel1-ef12800000000
W12-sel1-ef25600000000
W12-sel1-ef51200000000
W12-sel10-ef3200000000
W12-sel10-ef6400000000
W12-sel10-ef12800000000
W12-sel10-ef25600000000
W12-sel10-ef51200000000
W1300000000
W11-steady09.9 MiB016495.2 MiB3.7 GiB2,78766,097
W11039.4 MiB0652147.3 MiB5.8 GiB4,119102,476
Scheduler and memory CPU use, run-queue wait, migrations

Whether the engine was running while it ran. of wall is CPU over elapsed, waiting is runnable-but-not-scheduled, migrations checks the pinning claim, and switches gives voluntary over involuntary — the scheduler taking the core away against the engine choosing to sleep, which per unit of work is the cheapest signal of lock contention there is. A dash is not a zero: an index build's threads exit before the row does, and threads shows what fraction of the row the survivors account for.

Time spent waiting for a core

Lower is better; the bar between a pair is the gap.

Peak memory per row

Peak RSS while the row ran, from the engine's own process. Unlike the disk counters this is not refused across the two engines: residency decides where bytes live, and this is what the process held either way.

workloadstrawmANNQdrant
cpuof wallwaitingswitches vol/involmigrationsfaults min/majthreadscpuof wallwaitingswitches vol/involmigrationsfaults min/majthreads
W0-upload8.5 s138%3.4 ms4,121 / 36012,642 / 09 2%11.5 s198%857.4 ms16,417 / 1,8121,63246,470 / 111,01043 14%
W07.1 s68%13.1 ms168,401 / 700 / 0912.6 s103%3.7 ms561,307 / 19146377 / 037 0%
W10.6 s27%0.2 ms35,840 / 0056,822 / 09 99%7.1 s222%2,458.8 ms14,357 / 2,6263,17650,609 / 118,24353 41%
W256.5 s497%27.6 ms36,348 / 268062,863 / 09 1%71.3 s433%3,346.4 ms17,108 / 4,2683,68588,140 / 160,65044 6%
W326.7 s88%6.9 ms185,798 / 2306 / 0931.4 s103%3.7 ms560,703 / 51144363 / 039 0%
W434.1 s730%2.5 ms65,942 / 46065 / 0935.6 s771%82,694.5 ms57,535 / 41,72015,7942,042 / 076
W4-sat50117.4 s321%14.9 ms348,316 / 6308 / 09148.1 s405%119,429.2 ms1,210,934 / 124,899439,062314 / 068 89%
W4-sat70131.2 s501%3.7 ms278,186 / 54028 / 09153.0 s585%169,227.7 ms1,104,084 / 172,332452,191587 / 068
W4-sat90126.6 s621%15.2 ms253,342 / 86056 / 09139.2 s683%122,345.4 ms529,807 / 133,561178,0311,282 / 044 0%
W533.0 s702%2.1 ms9,862 / 5400 / 0926.4 s198%4.7 ms35,577 / 1801,355585 / 042 0%
W6-upload40.0 s385%17.0 ms7,384 / 1910128,700 / 09 1%34.5 s309%1,680.1 ms18,744 / 2,0472,722122,288 / 160,67742 0%
W614.9 s164%11.8 ms168,070 / 1100 / 0922.8 s202%12.3 ms502,120 / 3610,087430 / 044 0%
W6-ef326.2 s133%2.1 ms140,698 / 700 / 0913.1 s196%10.5 ms466,292 / 3210,58473 / 044
W6-ef649.0 s151%1.5 ms157,603 / 300 / 0916.6 s198%12.3 ms477,760 / 249,87387 / 044
W6-ef12815.3 s165%5.7 ms170,493 / 700 / 0922.6 s202%11.4 ms501,314 / 4310,05069 / 044
W6-ef25625.9 s176%1.3 ms177,766 / 1100 / 0933.6 s203%11.4 ms525,535 / 4912,715172 / 044
W6-ef51244.8 s186%5.8 ms182,326 / 3600 / 0954.6 s202%15.2 ms539,775 / 18517,661358 / 044
W7-upload17.1 s232%7.8 ms8,615 / 82064,793 / 09 2%31.1 s269%1,866.1 ms20,159 / 2,3662,983107,187 / 196,23442 0%
W714.1 s162%9.1 ms168,802 / 601 / 0921.7 s200%13.7 ms495,377 / 5310,273440 / 046 0%
W7-2bit-upload19.0 s226%7.4 ms8,250 / 99166,716 / 09 2%28.9 s250%1,449.4 ms20,120 / 2,1582,745100,202 / 180,73450 23%
W7-2bit14.9 s163%3.1 ms171,867 / 1400 / 0921.6 s200%11.9 ms491,664 / 4010,027380 / 050 0%
W7-1p5bit-upload19.4 s232%7.5 ms7,574 / 92166,708 / 09 2%33.2 s266%1,604.3 ms20,125 / 2,1522,756100,267 / 179,98251 0%
W7-1p5bit14.7 s162%3.4 ms170,077 / 900 / 0922.2 s201%15.1 ms495,584 / 3810,269463 / 050 0%
W8-upload342.9 s722%104.2 ms8,832 / 1,5830121,701 / 09 0%271.2 s563%1,586.7 ms106,277 / 11,3248,23892,993 / 126,80142 0%
W825.8 s177%1.4 ms176,141 / 1700 / 0935.1 s202%18.7 ms512,151 / 10212,396724 / 046
W14-1bit-upload57.8 s466%25.7 ms7,831 / 248067,661 / 09 1%33.9 s267%1,597.1 ms22,091 / 7,1693,84499,537 / 179,40851 0%
W14-1bit18.2 s167%11.2 ms171,064 / 1200 / 0922.1 s201%13.4 ms495,888 / 4310,473364 / 050 0%
W14-1p5bit-upload58.7 s473%27.5 ms7,897 / 284069,155 / 09 1%35.9 s274%1,853.0 ms21,159 / 3,9303,13299,394 / 163,22650 22%
W14-1p5bit19.5 s169%3.3 ms174,050 / 1100 / 0924.7 s202%9.1 ms515,129 / 4111,740511 / 050 0%
W14-2bit-upload58.1 s468%27.2 ms8,521 / 275071,668 / 09 1%36.5 s281%2,027.8 ms24,068 / 8,1124,012100,282 / 163,98850 20%
W14-2bit15.2 s165%2.5 ms170,628 / 1100 / 0921.9 s200%15.0 ms491,285 / 399,711154 / 050 0%
W14-4bit-upload58.8 s475%26.8 ms7,657 / 276079,674 / 09 1%37.2 s279%1,927.8 ms30,707 / 16,2304,040113,836 / 167,08651 0%
W14-4bit16.3 s165%3.9 ms172,180 / 800 / 0921.9 s201%15.2 ms494,084 / 549,940481 / 050 0%
W945.4 s704%4.3 ms3,835 / 7300 / 0949.8 s772%7,100.6 ms15,586 / 7,2235,75180 / 065
W10-ef3221.4 s584%3.3 ms162,478 / 1700 / 0916.1 s575%6,265.7 ms232,650 / 18,87665,474856 / 048 0%
W10-ef6418.9 s659%2.2 ms90,642 / 1600 / 0924.0 s592%14,521.1 ms274,069 / 34,041111,342524 / 072
W10-ef12833.3 s693%3.8 ms97,283 / 3200 / 0936.6 s602%26,270.7 ms295,037 / 35,762117,455214 / 066 0%
W10-ef25658.8 s704%4.7 ms104,254 / 6400 / 0958.6 s628%38,470.1 ms309,047 / 40,585143,422179 / 066
W10-ef512103.6 s708%8.4 ms111,925 / 13500 / 0996.4 s649%53,106.6 ms331,443 / 48,884169,278293 / 066
W12-upload56.5 s450%24.9 ms10,585 / 256063,572 / 09 1%85.5 s351%1,653.3 ms31,032 / 3,0923,459112,773 / 154,55551 0%
W12-sel111.2 s159%10.8 ms177,424 / 901 / 0929.7 s200%15.7 ms498,516 / 559,230621 / 056
W12-sel1071.2 s188%7.6 ms130,386 / 8700 / 0956.2 s202%13.6 ms537,860 / 13317,354381 / 056
W12-sel1-ef3211.0 s160%1.2 ms174,196 / 500 / 0917.2 s198%12.5 ms478,083 / 599,81036 / 056
W12-sel1-ef6411.0 s160%2.1 ms176,346 / 800 / 0922.2 s200%13.0 ms490,172 / 689,45162 / 056
W12-sel1-ef12811.1 s159%1.9 ms177,201 / 400 / 0930.2 s201%16.8 ms499,372 / 10610,052251 / 056
W12-sel1-ef25611.3 s159%2.2 ms176,052 / 400 / 0941.3 s201%22.1 ms504,917 / 16610,027218 / 056
W12-sel1-ef51211.7 s159%1.8 ms175,182 / 300 / 0959.5 s201%41.6 ms509,883 / 25611,588313 / 056
W12-sel10-ef3227.7 s178%3.0 ms178,558 / 1200 / 0925.5 s202%10.8 ms515,979 / 10312,232195 / 056
W12-sel10-ef6444.7 s185%2.1 ms179,629 / 2300 / 0936.5 s202%13.1 ms528,464 / 17814,006292 / 056
W12-sel10-ef12871.3 s188%8.4 ms133,979 / 8300 / 0956.2 s202%16.4 ms538,067 / 29417,649260 / 056
W12-sel10-ef25671.3 s188%5.8 ms133,831 / 7200 / 0998.6 s202%53.8 ms548,624 / 50519,972756 / 056
W12-sel10-ef51271.2 s188%4.7 ms134,101 / 7900 / 09162.6 s199%91.6 ms561,935 / 1,03221,3931,381 / 056
W135.6 s137%2.5 ms245,517 / 700 / 099.1 s409%922.2 ms382,586 / 5,038113,126622 / 067
W11-steady228.5 s685%3,877.6 ms641,262 / 1,3861623,682 / 09 98%277.1 s624%247,440.9 ms2,134,190 / 335,030894,806376,849 / 480,76674 95%
W11634.6 s689%18,357.3 ms1,669,390 / 7,1576089,971 / 09 97%720.3 s617%572,136.9 ms5,316,734 / 716,4052,221,296658,653 / 439,85374 95%
Waiting and stalls where the time went when the engine was not running

waiting is runnable and not scheduled, blocked on disk is not runnable at all, and pressure is PSI for the engine's cgroup (some = at least one task stalled, full = every runnable task).

Time stalled on I/O

Time every task in the engine's cgroup was stalled on I/O. Lower is better. 40 rows never stalled on I/O and are not drawn.

workloadstrawmANNQdrant
waitingblocked on diskcpu some/fullio some/fullmemory some/fullwaitingblocked on diskcpu some/fullio some/fullmemory some/full
W0-upload3.4 ms0 ms2 / 2 ms0 / 0 ms0 / 0 ms857.4 ms0 ms96 / 2 ms13 / 13 ms0 / 0 ms
W013.1 ms0 ms48 / 48 ms0 / 0 ms0 / 0 ms3.7 ms0 ms69 / 69 ms0 / 0 ms0 / 0 ms
W10.2 ms0 ms3 / 3 ms0 / 0 ms0 / 0 ms2,458.8 ms0 ms227 / 2 ms14 / 9 ms0 / 0 ms
W227.6 ms0 ms7 / 4 ms0 / 0 ms0 / 0 ms3,346.4 ms0 ms326 / 5 ms53 / 45 ms0 / 0 ms
W36.9 ms0 ms48 / 48 ms0 / 0 ms0 / 0 ms3.7 ms0 ms85 / 85 ms0 / 0 ms0 / 0 ms
W42.5 ms0 ms3 / 3 ms0 / 0 ms0 / 0 ms82,694.5 ms0 ms3,283 / 3 ms0 / 0 ms0 / 0 ms
W4-sat5014.9 ms0 ms86 / 86 ms0 / 0 ms0 / 0 ms119,429.2 ms0 ms8,875 / 109 ms0 / 0 ms0 / 0 ms
W4-sat703.7 ms0 ms55 / 55 ms0 / 0 ms0 / 0 ms169,227.7 ms0 ms11,072 / 75 ms0 / 0 ms0 / 0 ms
W4-sat9015.2 ms0 ms28 / 28 ms0 / 0 ms0 / 0 ms122,345.4 ms0 ms8,508 / 40 ms0 / 0 ms0 / 0 ms
W52.1 ms0 ms1 / 1 ms0 / 0 ms0 / 0 ms4.7 ms0 ms14 / 13 ms0 / 0 ms0 / 0 ms
W6-upload17.0 ms0 ms4 / 2 ms0 / 0 ms0 / 0 ms1,680.1 ms0 ms162 / 4 ms55 / 54 ms0 / 0 ms
W611.8 ms0 ms30 / 30 ms0 / 0 ms0 / 0 ms12.3 ms0 ms56 / 55 ms0 / 0 ms0 / 0 ms
W6-ef322.1 ms0 ms23 / 23 ms0 / 0 ms0 / 0 ms10.5 ms0 ms47 / 46 ms0 / 0 ms0 / 0 ms
W6-ef641.5 ms0 ms23 / 23 ms0 / 0 ms0 / 0 ms12.3 ms0 ms52 / 51 ms0 / 0 ms0 / 0 ms
W6-ef1285.7 ms0 ms30 / 30 ms0 / 0 ms0 / 0 ms11.4 ms0 ms55 / 55 ms0 / 0 ms0 / 0 ms
W6-ef2561.3 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms11.4 ms0 ms60 / 59 ms0 / 0 ms0 / 0 ms
W6-ef5125.8 ms0 ms36 / 36 ms0 / 0 ms0 / 0 ms15.2 ms0 ms69 / 68 ms0 / 0 ms0 / 0 ms
W7-upload7.8 ms0 ms2 / 2 ms0 / 0 ms0 / 0 ms1,866.1 ms0 ms227 / 5 ms37 / 34 ms0 / 0 ms
W79.1 ms0 ms30 / 30 ms0 / 0 ms0 / 0 ms13.7 ms0 ms57 / 56 ms0 / 0 ms0 / 0 ms
W7-2bit-upload7.4 ms0 ms2 / 2 ms0 / 0 ms0 / 0 ms1,449.4 ms0 ms172 / 5 ms49 / 48 ms0 / 0 ms
W7-2bit3.1 ms0 ms32 / 32 ms0 / 0 ms0 / 0 ms11.9 ms0 ms57 / 56 ms0 / 0 ms0 / 0 ms
W7-1p5bit-upload7.5 ms0 ms3 / 2 ms0 / 0 ms0 / 0 ms1,604.3 ms0 ms166 / 5 ms47 / 45 ms0 / 0 ms
W7-1p5bit3.4 ms0 ms29 / 29 ms0 / 0 ms0 / 0 ms15.1 ms0 ms58 / 57 ms0 / 0 ms0 / 0 ms
W8-upload104.2 ms0 ms15 / 4 ms0 / 0 ms0 / 0 ms1,586.7 ms0 ms342 / 35 ms51 / 46 ms0 / 0 ms
W81.4 ms0 ms32 / 32 ms0 / 0 ms0 / 0 ms18.7 ms0 ms70 / 69 ms0 / 0 ms0 / 0 ms
W14-1bit-upload25.7 ms0 ms5 / 2 ms0 / 0 ms0 / 0 ms1,597.1 ms0 ms221 / 6 ms70 / 67 ms0 / 0 ms
W14-1bit11.2 ms0 ms34 / 34 ms0 / 0 ms0 / 0 ms13.4 ms0 ms58 / 57 ms0 / 0 ms0 / 0 ms
W14-1p5bit-upload27.5 ms0 ms5 / 2 ms7 / 7 ms0 / 0 ms1,853.0 ms0 ms231 / 7 ms96 / 94 ms0 / 0 ms
W14-1p5bit3.3 ms0 ms33 / 33 ms0 / 0 ms0 / 0 ms9.1 ms0 ms57 / 56 ms0 / 0 ms0 / 0 ms
W14-2bit-upload27.2 ms0 ms5 / 2 ms1 / 1 ms0 / 0 ms2,027.8 ms0 ms260 / 7 ms58 / 54 ms0 / 0 ms
W14-2bit2.5 ms0 ms27 / 27 ms0 / 0 ms0 / 0 ms15.0 ms0 ms57 / 56 ms0 / 0 ms0 / 0 ms
W14-4bit-upload26.8 ms0 ms5 / 2 ms0 / 0 ms0 / 0 ms1,927.8 ms0 ms237 / 8 ms55 / 52 ms0 / 0 ms
W14-4bit3.9 ms0 ms32 / 32 ms0 / 0 ms0 / 0 ms15.2 ms0 ms56 / 55 ms0 / 0 ms0 / 0 ms
W94.3 ms0 ms1 / 1 ms0 / 0 ms0 / 0 ms7,100.6 ms0 ms823 / 11 ms0 / 0 ms0 / 0 ms
W10-ef323.3 ms0 ms21 / 21 ms0 / 0 ms0 / 0 ms6,265.7 ms0 ms715 / 16 ms0 / 0 ms0 / 0 ms
W10-ef642.2 ms0 ms11 / 11 ms0 / 0 ms0 / 0 ms14,521.1 ms0 ms1,351 / 18 ms0 / 0 ms0 / 0 ms
W10-ef1283.8 ms0 ms10 / 10 ms0 / 0 ms0 / 0 ms26,270.7 ms0 ms2,265 / 21 ms0 / 0 ms0 / 0 ms
W10-ef2564.7 ms0 ms9 / 9 ms0 / 0 ms0 / 0 ms38,470.1 ms0 ms3,489 / 24 ms0 / 0 ms0 / 0 ms
W10-ef5128.4 ms0 ms7 / 7 ms0 / 0 ms0 / 0 ms53,106.6 ms0 ms5,089 / 27 ms0 / 0 ms0 / 0 ms
W12-upload24.9 ms0 ms6 / 3 ms0 / 0 ms0 / 0 ms1,653.3 ms0 ms164 / 9 ms44 / 43 ms0 / 0 ms
W12-sel110.8 ms0 ms25 / 25 ms0 / 0 ms0 / 0 ms15.7 ms0 ms58 / 56 ms0 / 0 ms0 / 0 ms
W12-sel107.6 ms0 ms35 / 35 ms0 / 0 ms0 / 0 ms13.6 ms0 ms71 / 70 ms0 / 0 ms0 / 0 ms
W12-sel1-ef321.2 ms0 ms22 / 22 ms0 / 0 ms0 / 0 ms12.5 ms0 ms51 / 50 ms0 / 0 ms0 / 0 ms
W12-sel1-ef642.1 ms0 ms22 / 22 ms0 / 0 ms0 / 0 ms13.0 ms0 ms55 / 53 ms0 / 0 ms0 / 0 ms
W12-sel1-ef1281.9 ms0 ms24 / 24 ms0 / 0 ms0 / 0 ms16.8 ms0 ms58 / 57 ms0 / 0 ms0 / 0 ms
W12-sel1-ef2562.2 ms0 ms24 / 24 ms0 / 0 ms0 / 0 ms22.1 ms0 ms62 / 60 ms0 / 0 ms0 / 0 ms
W12-sel1-ef5121.8 ms0 ms24 / 24 ms0 / 0 ms0 / 0 ms41.6 ms0 ms71 / 67 ms0 / 0 ms0 / 0 ms
W12-sel10-ef323.0 ms0 ms34 / 34 ms0 / 0 ms0 / 0 ms10.8 ms0 ms57 / 56 ms0 / 0 ms0 / 0 ms
W12-sel10-ef642.1 ms0 ms34 / 34 ms0 / 0 ms0 / 0 ms13.1 ms0 ms62 / 61 ms0 / 0 ms0 / 0 ms
W12-sel10-ef1288.4 ms0 ms35 / 35 ms0 / 0 ms0 / 0 ms16.4 ms0 ms70 / 69 ms0 / 0 ms0 / 0 ms
W12-sel10-ef2565.8 ms0 ms35 / 35 ms0 / 0 ms0 / 0 ms53.8 ms0 ms112 / 107 ms0 / 0 ms0 / 0 ms
W12-sel10-ef5124.7 ms0 ms35 / 35 ms0 / 0 ms0 / 0 ms91.6 ms0 ms142 / 135 ms0 / 0 ms0 / 0 ms
W132.5 ms0 ms30 / 30 ms0 / 0 ms0 / 0 ms922.2 ms0 ms128 / 27 ms0 / 0 ms0 / 0 ms
W11-steady3,877.6 ms0 ms830 / 66 ms0 / 0 ms0 / 0 ms247,440.9 ms0 ms17,953 / 152 ms201 / 46 ms0 / 0 ms
W1118,357.3 ms0 ms3,736 / 176 ms10 / 10 ms0 / 0 ms572,136.9 ms0 ms44,599 / 386 ms348 / 82 ms0 / 0 ms

A dash under blocked on disk is not a zero: most hosts ship with kernel.task_delayacct=0, which reports the field as a permanent zero and would manufacture the strongest claim here — that the memory-resident engine never waits for a device — out of a sysctl. bench/setup.py apply turns it on.

Hardware counters cycles, DRAM loads, TLB walks, IPC per query

What the core did per query, from perf stat attached to the engine for the same window as the /proc counters. §5's cost model is a set of claims about instructions per cycle, DRAM traffic and TLB reach; these are those quantities, measured on the row the headline quotes rather than on a microbenchmark.

What one query cost, in cycles

Lower is better; the bar between a pair is the gap. Load rows are absent: they have no queries to divide by, and an absolute count under a per-query axis would be a different quantity wearing this one's label.

Demand loads from DRAM per query

Demand loads served from DRAM, times the cache line this host reports. Lower is better. Demand only: the hardware prefetcher's fills are a separate counter and are not in this number, so a row that streams — an exact scan above all — moved far more than this says. On the graph rows, where there is little for a prefetcher to predict, it is most of the traffic and is the outstanding-miss quantity §5.2 argues the engine is limited by.

TLB walks per query

Data TLB misses that reached a page walk, per query. Lower is better. §5.5 argues a memory-resident index needs hugepage care; this is what not taking it costs, and `--no-huge-pages` is the A/B arm that prices it.

Instructions per cycle

How well the core was fed while it ran. Not a score: a low IPC on a memory-bound row is what §5.2 predicts, and a high one on a row that does more work per query is not a win. Read it beside the two charts above.

Branch mispredictions per 1k instructions

A graph traversal is a chain of data-dependent branches and none of them is predictable from the last query, so this is the column that separates a slow kernel from an unpredictable walk. Per thousand instructions rather than per branch: `branches` costs a PMU counter that `ref-cycles` needs more, and MPKI is the comparable form anyway.

workloadstrawmANNQdrant
IPCGHzbranch MPKIcycles/querydemand DRAM/queryTLB walks/queryIPCGHzbranch MPKIcycles/querydemand DRAM/queryTLB walks/query
W0-upload1.952.011.005x nominal5.37---1.862.011.006x nominal4.49---
W01.502.011.006x nominal5.35233,4694.4 KiB582.91.432.011.006x nominal3.16374,5747.2 KiB725.9
W11.852.011.006x nominal0.84---1.852.011.005x nominal2.26---
W20.772.011.005x nominal3.78---0.932.011.006x nominal3.39---
W30.702.011.006x nominal5.69985,8861,953.5 KiB2,214.80.942.011.006x nominal3.261,128,5932,073.0 KiB3,147.9
W40.502.011.006x nominal4.871,356,6651,838.9 KiB1,682.00.692.011.006x nominal3.781,371,5561,717.6 KiB2,313.0
W4-sat500.602.011.006x nominal5.391,146,7371,895.8 KiB2,026.40.772.011.006x nominal3.841,360,2801,833.2 KiB2,864.4
W4-sat700.522.011.006x nominal5.321,294,2161,898.4 KiB1,918.80.752.011.007x nominal3.831,396,4711,786.9 KiB2,757.8
W4-sat900.542.011.007x nominal4.891,259,2131,873.3 KiB1,820.70.732.011.007x nominal3.771,333,5801,732.3 KiB2,486.2
W50.472.011.006x nominal5.161,325,9951,910.1 KiB1,671.30.772.011.006x nominal4.341,037,5012,140.9 KiB2,182.3
W6-upload1.212.011.007x nominal3.57---2.342.011.006x nominal2.84---
W61.482.011.006x nominal5.51547,931243.8 KiB2,262.41.732.011.006x nominal2.97797,572254.8 KiB2,579.5
W6-ef321.462.011.007x nominal3.27219,35260.4 KiB960.31.732.011.006x nominal2.17422,32168.4 KiB1,214.5
W6-ef641.502.011.006x nominal4.06325,096113.3 KiB1,406.81.712.011.006x nominal2.56559,582131.8 KiB1,762.6
W6-ef1281.442.011.007x nominal5.55562,388243.5 KiB2,265.71.742.011.007x nominal2.98793,797256.0 KiB2,583.4
W6-ef2561.482.011.007x nominal6.05973,275466.1 KiB3,698.41.762.011.006x nominal3.351,225,670497.2 KiB3,949.7
W6-ef5121.542.011.007x nominal6.311,731,137919.0 KiB6,511.61.742.011.006x nominal3.672,057,181975.3 KiB6,289.6
W7-upload1.812.011.005x nominal6.70---2.702.011.005x nominal3.77---
W71.332.011.006x nominal7.98517,567191.0 KiB1,523.31.752.011.006x nominal3.46756,751204.4 KiB1,992.1
W7-2bit-upload1.702.011.005x nominal6.34---2.502.011.006x nominal3.83---
W7-2bit1.312.011.005x nominal7.87546,372196.8 KiB1,860.81.682.011.006x nominal3.42754,912207.4 KiB2,304.4
W7-1p5bit-upload1.732.011.006x nominal6.19---2.612.011.006x nominal3.77---
W7-1p5bit1.342.011.006x nominal7.67543,907196.9 KiB1,885.31.722.011.006x nominal3.45774,544215.6 KiB2,237.1
W8-upload5.982.011.006x nominal0.54---5.822.011.006x nominal0.25---
W83.402.011.006x nominal1.61971,825199.6 KiB1,976.02.862.011.007x nominal1.161,276,102204.9 KiB2,477.8
W14-1bit-upload0.882.011.005x nominal3.30---2.702.011.005x nominal2.62---
W14-1bit1.602.011.006x nominal4.69675,022189.9 KiB1,727.91.662.011.006x nominal3.03768,426208.8 KiB1,913.3
W14-1p5bit-upload0.922.011.006x nominal3.13---2.672.011.005x nominal2.27---
W14-1p5bit1.732.011.006x nominal4.16726,348195.2 KiB1,917.31.592.011.005x nominal2.78874,531258.7 KiB2,038.6
W14-2bit-upload0.882.011.005x nominal3.33---2.782.011.006x nominal1.94---
W14-2bit1.702.011.006x nominal4.96559,282199.4 KiB1,965.01.662.011.006x nominal2.99762,046216.7 KiB2,121.7
W14-4bit-upload0.902.011.005x nominal3.35---2.622.011.006x nominal2.03---
W14-4bit1.642.011.006x nominal4.94600,297208.3 KiB2,215.81.662.011.006x nominal2.95764,630234.3 KiB2,331.9
W90.452.011.005x nominal0.2045,430,23338,805.8 KiB13,036.90.552.011.006x nominal0.2349,311,81814,364.0 KiB50,750.0
W10-ef320.692.011.007x nominal3.46409,203629.3 KiB633.40.962.011.007x nominal2.72580,475680.1 KiB1,027.0
W10-ef640.572.011.007x nominal4.19738,6511,059.4 KiB990.20.852.011.006x nominal3.29869,6711,094.7 KiB1,642.2
W10-ef1280.522.011.007x nominal4.871,316,1021,871.8 KiB1,720.80.772.011.006x nominal3.801,352,1641,800.8 KiB2,622.0
W10-ef2560.502.011.007x nominal5.412,340,5423,351.8 KiB3,254.30.732.011.007x nominal4.192,193,4033,059.3 KiB4,225.3
W10-ef5120.512.011.007x nominal5.774,131,2185,975.3 KiB6,506.70.712.011.006x nominal4.493,689,4175,217.7 KiB6,967.8
W12-upload0.772.011.006x nominal3.78---1.182.011.006x nominal3.26---
W12-sel10.902.011.007x nominal3.54409,3211,385.0 KiB1,351.31.152.011.006x nominal4.161,084,0371,061.7 KiB2,369.7
W12-sel100.832.011.007x nominal1.162,809,4339,484.6 KiB8,566.11.162.011.007x nominal2.752,120,9992,510.2 KiB4,006.1
W12-sel1-ef320.922.011.006x nominal3.43399,8981,358.8 KiB1,350.61.182.011.006x nominal2.92583,982526.1 KiB1,542.1
W12-sel1-ef640.922.011.006x nominal3.43401,3491,360.8 KiB1,345.21.152.011.006x nominal3.52779,509770.0 KiB1,943.1
W12-sel1-ef1280.922.011.005x nominal3.43402,4941,363.0 KiB1,346.31.142.011.006x nominal4.201,096,1581,066.9 KiB2,369.2
W12-sel1-ef2560.892.011.006x nominal3.49411,2111,360.0 KiB1,363.01.212.011.006x nominal4.811,538,1101,356.5 KiB2,796.1
W12-sel1-ef5120.872.011.006x nominal3.59425,1181,376.1 KiB1,368.71.302.011.007x nominal5.302,262,1901,630.4 KiB3,189.2
W12-sel10-ef321.282.011.007x nominal3.981,047,5461,199.4 KiB2,529.91.192.011.007x nominal2.35904,771898.0 KiB2,036.0
W12-sel10-ef641.272.011.007x nominal4.271,727,6932,023.7 KiB3,981.31.162.011.006x nominal2.531,339,2191,495.5 KiB2,789.9
W12-sel10-ef1280.832.011.006x nominal1.152,812,0219,484.7 KiB8,566.41.162.011.007x nominal2.752,125,3202,509.3 KiB4,004.9
W12-sel10-ef2560.832.011.006x nominal1.152,809,8229,472.2 KiB8,564.91.082.011.007x nominal3.273,780,5874,341.7 KiB5,854.9
W12-sel10-ef5120.832.011.007x nominal1.162,812,6249,476.0 KiB8,562.51.112.011.006x nominal3.466,315,3946,872.6 KiB8,444.3
W131.362.011.007x nominal1.3142,5631.0 KiB14.61.452.011.006x nominal1.50140,1291.6 KiB40.0
W11-steady0.522.011.006x nominal4.831,363,6521,951.3 KiB1,792.20.852.011.006x nominal3.481,532,6151,896.2 KiB2,737.5
W110.512.011.007x nominal4.731,451,3742,074.0 KiB1,927.60.822.011.007x nominal3.401,526,5401,952.6 KiB2,803.9

Nothing here is scaled. Where perf had to multiplex the group, the values it prints are extrapolations from the fraction of the row each counter was on, and they are withheld rather than shown — the same rule the scheduler table applies to a thread that exited. An event this host does not implement is likewise blank, never zero.

The two instruments disagree on these rows. Qdrant: perf and /proc disagree on minor faults by up to 35% on W0-upload, W1, W2, W6-upload, W7-upload, W7-2bit-upload, W7-1p5bit-upload, W8-upload, W14-1bit-upload, W14-1p5bit-upload, W14-2bit-upload, W14-4bit-upload, W12-upload, W11-steady, W11; Qdrant: perf and /proc disagree on context switches by up to 70% on W1, W7-upload, W7-2bit-upload, W7-1p5bit-upload, W8-upload, W14-1bit-upload, W14-1p5bit-upload, W14-2bit-upload, W14-4bit-upload. perf keeps an exiting thread's counts and the /proc sums do not, so a gap here is usually the same lost-thread effect the scheduler table reports as coverage.

Collection config what each engine says it built

Read back from the engine rather than taken from the request. --segments 1 sets Qdrant's default_segment_number, which its optimizer treats as a target, and the segment count is the largest confound in the ef comparison. strawmANN held 1 segment; Qdrant held 2 segments.

Where the engines disagree a cell reads strawmANN / Qdrant, and is marked.

collectionsegmentspopulated segmentsrequested segmentspointsindexed vectorsvector sizehnsw mhnsw ef_constructquantizationshardsvector residency
bench01 / 2- / 1- / 1100,000100,000416100-1cached
bench11 / 5- / 5- / 1100,0000 / 19,10051216100-1cached
bench21 / 2- / 1- / 1100,000100,00051216100-1cached
bench61 / 2- / 1- / 1100,000100,00051216100scalar1cached
bench71 / 2- / 1- / 1100,000100,00051216100binary1cached
bench7b21 / 2- / 1- / 1100,000100,00051216100binary1cached
bench7b151 / 2- / 1- / 1100,000100,00051216100binary1cached
bench81 / 2- / 1- / 1100,000100,00051216100product1cached
bench14t11 / 2- / 1- / 1100,000100,00051216100turboquant1cached
bench14t151 / 2- / 1- / 1100,000100,00051216100turboquant1cached
bench14t21 / 2- / 1- / 1100,000100,00051216100turboquant1cached
bench14t41 / 2- / 1- / 1100,000100,00051216100turboquant1cached
bench121 / 2- / 1- / 1100,000100,00051216100-1cached

After the mutating rows bench2 held strawmANN 125,000 (+25,000), 125,000 of them indexed; Qdrant 125,000 (+25,000), 125,000 of them indexed. The table above is the state the throughput, latency and recall rows searched; this is what W11 left, and is that row's subject rather than theirs.

bench1 was read back straight after its row and dropped, with no settle between, so an engine still building it reports a snapshot taken mid-ingest. Its cells are not marked as a disagreement.

Per workload every metric of every row, engine beside engine

Metric down, engine across, so a comparison is two adjacent cells. Metrics a row did not measure are dropped rather than shown empty.

W0-upload transport floor: load

upload · collection=bench0 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock6.2 s5.8 s0.94x
of which upload0.1 s0.5 s4.20x
of which index wait4.0 s3.0 s0.75x
time to Green4.0 s3.0 s (at bfb's polling floor)0.75x
cpu8.5 s11.5 s1.35x
cpu, % of wall138%198%
waiting for a core3.4 ms857.4 ms252.69x
migrations01,632
faults min/maj12,642 / 046,470 / 111,010
peak RSS144.6 MiB688.1 MiB4.76x
disk written6.1 MiB35.0 MiB
disk read010.2 MiB

W0 d=4 floor: graph traversal with the distance taken out 1.18x

closed-loop · collection=bench0 · client -p 1 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second4,8244,0751.18x
client p50205 µs242 µs1.18x
client p99250 µs300 µs1.20x
client p99.9305 µs364 µs1.19x
server p50112 µs133 µs1.19x
server p99122 µs169 µs1.38x
wall clock10.4 s12.3 s1.18x
queries sent50,00050,000
cpu7.1 s12.6 s1.78x
cpu, % of wall68%103%
waiting for a core13.1 ms3.7 ms0.28x
migrations0146
faults min/maj0 / 0377 / 0
peak RSS144.6 MiB688.1 MiB4.76x
disk written00
disk read00
1.18x is 1.60x less work per query x 0.66x cores busy during the row (0.66 against 1.00) x 1.00x clock x 1.12x counted on-CPU share: the middle term is occupancy, not search speed d=4 floor: this measures graph traversal and heaps with the distance taken out, not request plumbing (findings 28: 76% of the server's CPU was HNSW search, 24% kernel; the RPC path is a small remainder)

W1 ingest throughput

upload · collection=bench1 · client -p 8 -t 8 -c 1 (-c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock2.3 s3.2 s1.37x
of which upload0.3 s1.1 s3.86x
cpu0.6 s7.1 s11.25x
cpu, % of wall27%222%
waiting for a core0.2 ms2,458.8 ms13,114.14x
migrations03,176
faults min/maj56,822 / 050,609 / 118,243
peak RSS321.8 MiB1,013.4 MiB3.15x
disk written195.4 MiB436.5 MiB
disk read0768.0 KiB

W2 index build time

upload · collection=bench2 · client -p 8 -t 8 -c 1 (-c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock11.4 s16.4 s1.45x
of which upload0.3 s1.2 s4.16x
of which index wait9.0 s13.0 s1.44x
time to Green9.0 s13.0 s1.44x
cpu56.5 s71.3 s1.26x
cpu, % of wall497%433%
waiting for a core27.6 ms3,346.4 ms121.08x
migrations03,685
faults min/maj62,863 / 088,140 / 160,650
peak RSS346.5 MiB1.9 GiB5.48x
disk written195.4 MiB843.1 MiB
disk read011.9 MiB

W3 search, fp32, single query parity

ef=128 · closed-loop · collection=bench2 · client -p 1 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,6461,637parity
recall@100.99400.9906~equal
client p50574 µs611 µs1.06x
client p99966 µs790 µs0.82x
client p99.91.07 ms860 µs0.80x
server p50471 µs511 µs1.08x
server p99852 µs683 µs0.80x
wall clock30.4 s30.6 s~equal
queries sent50,00050,000
cpu26.7 s31.4 s1.18x
cpu, % of wall88%103%
waiting for a core6.9 ms3.7 ms0.54x
migrations0144
faults min/maj6 / 0363 / 0
peak RSS346.5 MiB1.9 GiB5.48x
disk written00
disk read00
within the ±2.3% band this dataset's noise floor puts on W3: no measured difference, not a small one

W4 search, saturating (closed loop) parity

ef=128 · closed-loop · collection=bench2 · client -p 64 -t 16 -c 2 · n=50,000
metricstrawmANNQdrant
queries/second10,80010,924parity
recall@100.99400.9906~equal
client p505.90 ms5.36 ms0.91x
client p996.54 ms14.89 ms2.28x
client p99.919.21 ms26.91 ms1.40x
server p505.80 ms4.35 ms0.75x
server p996.42 ms11.31 ms1.76x
wall clock4.7 s4.6 s~equal
queries sent50,00050,000
cpu34.1 s35.6 s1.04x
cpu, % of wall730%771%
waiting for a core2.5 ms82,694.5 ms32,895.25x
migrations015,794
faults min/maj65 / 02,042 / 0
peak RSS346.5 MiB1.9 GiB5.48x
disk written00
disk read00
within the ±2.0% band this dataset's noise floor puts on W4: no measured difference, not a small one

W4-sat50 search, fixed rate at 50% of saturation (open loop) offered

ef=128 · open-loop · offered=5,489/s (50% of saturation, pinned reference 10,978 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second5,4895,489offered
recall@100.99400.9906~equal
client p50696 µs856 µs1.23x
client p991.34 ms1.95 ms1.45x
client p99.91.59 ms2.39 ms1.51x
server p50583 µs694 µs1.19x
server p991.19 ms1.74 ms1.46x
wall clock36.6 s36.6 s~equal
queries sent200,000200,000
cpu117.4 s148.1 s1.26x
cpu, % of wall321%405%
waiting for a core14.9 ms119,429.2 ms8,035.02x
migrations0439,062
faults min/maj8 / 0314 / 0
peak RSS346.5 MiB1.9 GiB5.48x
disk written00
disk read00

W4-sat70 search, fixed rate at 70% of saturation (open loop) offered

ef=128 · open-loop · offered=7,685/s (70% of saturation, pinned reference 10,978 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second7,6857,685offered
recall@100.99400.9906~equal
client p50861 µs942 µs1.09x
client p991.89 ms2.20 ms1.16x
client p99.92.21 ms3.30 ms1.50x
server p50742 µs757 µs1.02x
server p991.70 ms1.89 ms1.11x
wall clock26.2 s26.2 s~equal
queries sent200,000200,000
cpu131.2 s153.0 s1.17x
cpu, % of wall501%585%
waiting for a core3.7 ms169,227.7 ms45,936.04x
migrations0452,191
faults min/maj28 / 0587 / 0
peak RSS346.5 MiB1.9 GiB5.48x
disk written00
disk read00

W4-sat90 search, fixed rate at 90% of saturation (open loop) offered

ef=128 · open-loop · offered=9,880/s (90% of saturation, pinned reference 10,978 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second9,8799,879offered
recall@100.99400.9906~equal
client p50913 µs1.33 ms1.46x
client p992.14 ms4.91 ms2.30x
client p99.94.15 ms8.14 ms1.96x
server p50790 µs976 µs1.23x
server p991.93 ms3.60 ms1.86x
wall clock20.4 s20.4 s~equal
queries sent200,000200,000
cpu126.6 s139.2 s1.10x
cpu, % of wall621%683%
waiting for a core15.2 ms122,345.4 ms8,051.04x
migrations0178,031
faults min/maj56 / 01,282 / 0
peak RSS346.5 MiB1.9 GiB5.48x
disk written00
disk read00

W5 search batched (16 distinct dataset queries per request) 2.83x

ef=128 · closed-loop · collection=bench2 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second10,6433,7542.83x
recall@100.99400.9906~equal
client p503.00 ms8.50 ms2.83x
client p993.52 ms9.40 ms2.67x
client p99.93.69 ms9.87 ms2.68x
server p502.81 ms8.00 ms2.85x
server p993.32 ms8.86 ms2.67x
wall clock4.7 s13.3 s2.83x
queries sent50,00050,000
cpu33.0 s26.4 s0.80x
cpu, % of wall702%198%
waiting for a core2.1 ms4.7 ms2.22x
migrations01,355
faults min/maj0 / 0585 / 0
peak RSS346.5 MiB1.9 GiB5.48x
disk written00
disk read00
per-batch latency (16 queries/request); queries: dataset, random-sample 2.83x is 0.78x less work per query x 3.58x cores busy during the row (7.04 against 1.97) x 1.00x clock x 1.01x counted on-CPU share: the middle term is occupancy, not search speed batched 16, so qps and rps differ by 16x; the ratio uses qps

W6-upload scalar quantization: load

upload · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock10.4 s11.2 s1.07x
of which upload0.3 s1.0 s3.16x
of which index wait8.0 s8.0 s~equal
time to Green8.0 s8.0 s~equal
cpu40.0 s34.5 s0.86x
cpu, % of wall385%309%
waiting for a core17.0 ms1,680.1 ms98.65x
migrations02,722
faults min/maj128,700 / 0122,288 / 160,677
peak RSS877.6 MiB2.5 GiB2.96x
disk written195.4 MiB992.2 MiB
disk read012.7 MiB

W6 quantized: scalar 1.24x

ef=128 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second5,5074,4431.24x
recall@100.99370.9908~equal
client p50355 µs447 µs1.26x
client p99538 µs534 µs~equal
client p99.9660 µs632 µs0.96x
server p50264 µs345 µs1.30x
server p99427 µs412 µs0.96x
wall clock9.1 s11.3 s1.24x
queries sent50,00050,000
cpu14.9 s22.8 s1.53x
cpu, % of wall164%202%
waiting for a core11.8 ms12.3 ms1.04x
migrations010,087
faults min/maj0 / 0430 / 0
peak RSS877.6 MiB2.5 GiB2.96x
disk written00
disk read00

W6-ef32 SQ8 recall control, ef=32 (latency only)

ef=32 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=3.2 · rescore=true
metricstrawmANNQdrant
queries/second10,8477,502-
recall@100.96240.9504
client p50180 µs260 µs
client p99247 µs344 µs
client p99.9280 µs389 µs
server p5091 µs159 µs
server p99140 µs220 µs
wall clock4.6 s6.7 s
queries sent50,00050,000
cpu6.2 s13.1 s
cpu, % of wall133%196%
waiting for a core2.1 ms10.5 ms
migrations010,584
faults min/maj0 / 073 / 0
peak RSS877.6 MiB2.5 GiB
disk written00
disk read00
recall unequal: 0.9624 vs 0.9504; §7.4 compares at equal recall; the rescore pools match, so the encoders differ

W6-ef64 SQ8 recall control, ef=64 (latency only) 1.41x

ef=64 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=6.4 · rescore=true
metricstrawmANNQdrant
queries/second8,4525,9881.41x
recall@100.98500.9783~equal
client p50229 µs327 µs1.43x
client p99324 µs437 µs1.35x
client p99.9369 µs491 µs1.33x
server p50142 µs226 µs1.59x
server p99218 µs317 µs1.45x
wall clock6.0 s8.4 s1.41x
queries sent50,00050,000
cpu9.0 s16.6 s1.85x
cpu, % of wall151%198%
waiting for a core1.5 ms12.3 ms8.09x
migrations09,873
faults min/maj0 / 087 / 0
peak RSS877.6 MiB2.5 GiB2.96x
disk written00
disk read00
1.41x is 1.72x less work per query x 0.77x cores busy during the row (1.51 against 1.97) x 1.00x clock x 1.07x counted on-CPU share: the middle term is occupancy, not search speed

W6-ef128 SQ8 recall control, ef=128 (latency only) 1.21x

ef=128 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second5,4024,4691.21x
recall@100.99370.9908~equal
client p50360 µs445 µs1.24x
client p99544 µs531 µs0.98x
client p99.9605 µs618 µs1.02x
server p50269 µs343 µs1.28x
server p99435 µs410 µs0.94x
wall clock9.3 s11.2 s1.21x
queries sent50,00050,000
cpu15.3 s22.6 s1.48x
cpu, % of wall165%202%
waiting for a core5.7 ms11.4 ms2.00x
migrations010,050
faults min/maj0 / 069 / 0
peak RSS877.6 MiB2.5 GiB2.96x
disk written00
disk read00

W6-ef256 SQ8 recall control, ef=256 (latency only) 1.13x

ef=256 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=25.6 · rescore=true
metricstrawmANNQdrant
queries/second3,4103,0251.13x
recall@100.99750.9958~equal
client p50576 µs662 µs1.15x
client p99868 µs779 µs0.90x
client p99.9950 µs842 µs0.89x
server p50483 µs559 µs1.16x
server p99757 µs663 µs0.88x
wall clock14.7 s16.6 s1.13x
queries sent50,00050,000
cpu25.9 s33.6 s1.29x
cpu, % of wall176%203%
waiting for a core1.3 ms11.4 ms8.80x
migrations012,715
faults min/maj0 / 0172 / 0
peak RSS877.6 MiB2.5 GiB2.96x
disk written00
disk read00

W6-ef512 SQ8 recall control, ef=512 (latency only)

ef=512 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=51.2 · rescore=true
metricstrawmANNQdrant
queries/second2,0721,858-
recall@100.99890.9980
client p50964 µs1.08 ms
client p991.31 ms1.28 ms
client p99.91.53 ms1.36 ms
server p50868 µs975 µs
server p991.20 ms1.16 ms
wall clock24.2 s27.0 s
queries sent50,00050,000
cpu44.8 s54.6 s
cpu, % of wall186%202%
waiting for a core5.8 ms15.2 ms
migrations017,661
faults min/maj0 / 0358 / 0
peak RSS877.6 MiB2.5 GiB
disk written00
disk read00
Qdrant pass 3 -21% against two passes that agree: one pass set the spread on this row, so a band built from it is not a noise band

W7-upload binary quantization: load

upload · collection=bench7 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock7.4 s11.5 s1.56x
of which upload0.3 s1.4 s4.15x
of which index wait5.0 s8.0 s1.60x
time to Green5.0 s8.0 s1.60x
cpu17.1 s31.1 s1.82x
cpu, % of wall232%269%
waiting for a core7.8 ms1,866.1 ms238.27x
migrations02,983
faults min/maj64,793 / 0107,187 / 196,234
peak RSS974.0 MiB3.1 GiB3.22x
disk written195.4 MiB973.9 MiB
disk read018.7 MiB

W7 quantized: binary + oversampling 1.25x

ef=128 · closed-loop · collection=bench7 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second5,7694,6161.25x
recall@100.91500.9162~equal
client p50341 µs423 µs1.24x
client p99469 µs550 µs1.17x
client p99.9629 µs607 µs0.97x
server p50249 µs319 µs1.28x
server p99360 µs424 µs1.18x
wall clock8.7 s10.9 s1.25x
queries sent50,00050,000
cpu14.1 s21.7 s1.54x
cpu, % of wall162%200%
waiting for a core9.1 ms13.7 ms1.51x
migrations010,273
faults min/maj1 / 0440 / 0
peak RSS974.0 MiB3.1 GiB3.22x
disk written00
disk read00

W7-2bit-upload binary quantization, 2 bits: load

upload · collection=bench7b2 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock8.4 s11.6 s1.38x
of which upload0.3 s1.3 s3.90x
of which index wait6.0 s8.0 s1.33x
time to Green6.0 s8.0 s1.33x
cpu19.0 s28.9 s1.52x
cpu, % of wall226%250%
waiting for a core7.4 ms1,449.4 ms196.07x
migrations12,7452,745.00x
faults min/maj66,716 / 0100,202 / 180,734
peak RSS1.5 GiB3.1 GiB2.03x
disk written195.4 MiB902.0 MiB
disk read014.1 MiB

W7-2bit quantized: binary quantization, 2 bits + oversampling 1.18x

ef=128 · closed-loop · collection=bench7b2 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second5,4654,6361.18x
recall@100.96960.9674~equal
client p50361 µs420 µs1.16x
client p99474 µs548 µs1.15x
client p99.9537 µs609 µs1.13x
server p50266 µs317 µs1.19x
server p99364 µs423 µs1.16x
wall clock9.2 s10.8 s1.18x
queries sent50,00050,000
cpu14.9 s21.6 s1.45x
cpu, % of wall163%200%
waiting for a core3.1 ms11.9 ms3.77x
migrations010,027
faults min/maj0 / 0380 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00

W7-1p5bit-upload binary quantization, 1.5 bits: load

upload · collection=bench7b15 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock8.4 s12.5 s1.49x
of which upload0.3 s1.3 s4.14x
of which index wait6.0 s9.0 s1.50x
time to Green6.0 s9.0 s1.50x
cpu19.4 s33.2 s1.71x
cpu, % of wall232%266%
waiting for a core7.5 ms1,604.3 ms212.53x
migrations12,7562,756.00x
faults min/maj66,708 / 0100,267 / 179,982
peak RSS1.5 GiB3.1 GiB2.03x
disk written195.4 MiB903.3 MiB
disk read015.6 MiB

W7-1p5bit quantized: binary quantization, 1.5 bits + oversampling 1.22x

ef=128 · closed-loop · collection=bench7b15 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second5,5254,5381.22x
recall@100.94750.9477~equal
client p50357 µs431 µs1.21x
client p99477 µs563 µs1.18x
client p99.9541 µs621 µs1.15x
server p50264 µs328 µs1.24x
server p99367 µs437 µs1.19x
wall clock9.1 s11.1 s1.22x
queries sent50,00050,000
cpu14.7 s22.2 s1.51x
cpu, % of wall162%201%
waiting for a core3.4 ms15.1 ms4.38x
migrations010,269
faults min/maj0 / 0463 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00

W8-upload PQ: load

upload · collection=bench8 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock47.5 s48.2 s~equal
of which upload0.3 s1.0 s3.04x
of which index wait45.2 s45.1 s~equal
time to Green45.2 s45.1 s~equal
cpu342.9 s271.2 s0.79x
cpu, % of wall722%563%
waiting for a core104.2 ms1,586.7 ms15.23x
migrations08,238
faults min/maj121,701 / 092,993 / 126,801
peak RSS1.4 GiB3.1 GiB2.15x
disk written195.4 MiB814.3 MiB
disk read011.6 MiB

W8 quantized: PQ 1.19x

ef=128 · closed-loop · collection=bench8 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second3,4402,8881.19x
recall@100.99270.9863~equal
client p50580 µs692 µs1.19x
client p99726 µs816 µs1.12x
client p99.9780 µs897 µs1.15x
server p50483 µs576 µs1.19x
server p99616 µs679 µs1.10x
wall clock14.6 s17.4 s1.19x
queries sent50,00050,000
cpu25.8 s35.1 s1.36x
cpu, % of wall177%202%
waiting for a core1.4 ms18.7 ms13.20x
migrations012,396
faults min/maj0 / 0724 / 0
peak RSS1.4 GiB3.1 GiB2.15x
disk written00
disk read00

W14-1bit-upload TurboQuant, 1 bit: load

upload · collection=bench14t1 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock12.4 s12.7 s1.02x
of which upload0.3 s1.5 s4.61x
of which index wait10.0 s9.0 s0.90x
time to Green10.0 s9.0 s0.90x
cpu57.8 s33.9 s0.59x
cpu, % of wall466%267%
waiting for a core25.7 ms1,597.1 ms62.08x
migrations03,844
faults min/maj67,661 / 099,537 / 179,408
peak RSS1.5 GiB3.1 GiB2.03x
disk written195.4 MiB898.4 MiB
disk read016.1 MiB

W14-1bit quantized: TurboQuant, 1 bit + oversampling parity

ef=128 · closed-loop · collection=bench14t1 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second4,6124,570parity
recall@100.98950.9880~equal
client p50433 µs429 µs~equal
client p99538 µs554 µs1.03x
client p99.9604 µs610 µs~equal
server p50330 µs326 µs~equal
server p99425 µs426 µs~equal
wall clock10.9 s11.0 s~equal
queries sent50,00050,000
cpu18.2 s22.1 s1.21x
cpu, % of wall167%201%
waiting for a core11.2 ms13.4 ms1.19x
migrations010,473
faults min/maj0 / 0364 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00
within the ±2.0% band this dataset's noise floor puts on W14-1bit: no measured difference, not a small one

W14-1p5bit-upload TurboQuant, 1.5 bits: load

upload · collection=bench14t15 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock12.4 s13.1 s1.06x
of which upload0.3 s1.9 s6.05x
of which index wait10.0 s9.0 s0.90x
time to Green10.0 s9.0 s0.90x
cpu58.7 s35.9 s0.61x
cpu, % of wall473%274%
waiting for a core27.5 ms1,853.0 ms67.32x
migrations03,132
faults min/maj69,155 / 099,394 / 163,226
peak RSS1.5 GiB3.1 GiB2.03x
disk written195.4 MiB880.6 MiB
disk read015.2 MiB

W14-1p5bit quantized: TurboQuant, 1.5 bits + oversampling 1.07x

ef=128 · closed-loop · collection=bench14t15 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second4,3564,0901.07x
recall@100.99220.9900~equal
client p50458 µs484 µs1.06x
client p99576 µs655 µs1.14x
client p99.9631 µs726 µs1.15x
server p50357 µs381 µs1.07x
server p99465 µs536 µs1.15x
wall clock11.5 s12.3 s1.06x
queries sent50,00050,000
cpu19.5 s24.7 s1.27x
cpu, % of wall169%202%
waiting for a core3.3 ms9.1 ms2.74x
migrations011,740
faults min/maj0 / 0511 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00

W14-2bit-upload TurboQuant, 2 bits: load

upload · collection=bench14t2 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock12.4 s13.0 s1.05x
of which upload0.3 s1.8 s5.58x
of which index wait10.0 s9.0 s0.90x
time to Green10.0 s9.0 s0.90x
cpu58.1 s36.5 s0.63x
cpu, % of wall468%281%
waiting for a core27.2 ms2,027.8 ms74.53x
migrations04,012
faults min/maj71,668 / 0100,282 / 163,988
peak RSS1.5 GiB3.1 GiB2.03x
disk written195.4 MiB889.9 MiB
disk read012.8 MiB

W14-2bit quantized: TurboQuant, 2 bits + oversampling 1.19x

ef=128 · closed-loop · collection=bench14t2 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second5,4644,5961.19x
recall@100.99350.9900~equal
client p50362 µs425 µs1.17x
client p99476 µs551 µs1.16x
client p99.9519 µs607 µs1.17x
server p50270 µs321 µs1.19x
server p99368 µs424 µs1.15x
wall clock9.2 s10.9 s1.19x
queries sent50,00050,000
cpu15.2 s21.9 s1.44x
cpu, % of wall165%200%
waiting for a core2.5 ms15.0 ms6.08x
migrations09,711
faults min/maj0 / 0154 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00

W14-4bit-upload TurboQuant, 4 bits: load

upload · collection=bench14t4 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock12.4 s13.3 s1.07x
of which upload0.3 s2.2 s6.73x
of which index wait10.0 s9.0 s0.90x
time to Green10.0 s9.0 s0.90x
cpu58.8 s37.2 s0.63x
cpu, % of wall475%279%
waiting for a core26.8 ms1,927.8 ms71.83x
migrations04,040
faults min/maj79,674 / 0113,836 / 167,086
peak RSS1.5 GiB3.1 GiB2.03x
disk written195.4 MiB912.4 MiB
disk read011.7 MiB

W14-4bit quantized: TurboQuant, 4 bits + oversampling 1.11x

ef=128 · closed-loop · collection=bench14t4 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second5,0864,5961.11x
recall@100.99370.9904~equal
client p50387 µs429 µs1.11x
client p99516 µs557 µs1.08x
client p99.9572 µs623 µs1.09x
server p50293 µs327 µs1.11x
server p99404 µs436 µs1.08x
wall clock9.9 s10.9 s1.11x
queries sent50,00050,000
cpu16.3 s21.9 s1.34x
cpu, % of wall165%201%
waiting for a core3.9 ms15.2 ms3.89x
migrations09,940
faults min/maj0 / 0481 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00

W9 exact / brute force parity

exact · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=2,000
metricstrawmANNQdrant
queries/second311311parity
client p5025.51 ms26.05 ms1.02x
client p9935.49 ms31.80 ms0.90x
client p99.939.53 ms36.92 ms0.93x
server p5025.08 ms25.39 ms~equal
server p9935.06 ms31.03 ms0.89x
wall clock6.4 s6.4 s~equal
queries sent2,0002,000
cpu45.4 s49.8 s1.10x
cpu, % of wall704%772%
waiting for a core4.3 ms7,100.6 ms1,662.09x
migrations05,751
faults min/maj0 / 080 / 0
peak RSS1.4 GiB3.1 GiB2.15x
disk written00
disk read00
within the ±12.2% band this dataset's noise floor puts on W9: no measured difference, not a small one brute force over the whole collection, no index involved

W10-ef32 recall control, ef=32 (latency only)

ef=32 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=100,000
metricstrawmANNQdrant
queries/second27,90118,065-
recall@100.96300.9513
client p50279 µs411 µs
client p99403 µs801 µs
client p99.9499 µs1.11 ms
server p50185 µs274 µs
server p99293 µs577 µs
wall clock3.7 s2.8 s
queries sent100,00050,000
cpu21.4 s16.1 s
cpu, % of wall584%575%
waiting for a core3.3 ms6,265.7 ms
migrations065,474
faults min/maj0 / 0856 / 0
peak RSS1.4 GiB3.1 GiB
disk written00
disk read00
strawmANN: -n 100000 (2x the table's 50000) recall unequal: 0.9630 vs 0.9513; §7.4 compares at equal recall

W10-ef64 recall control, ef=64 (latency only) 1.42x

ef=64 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second17,66712,4381.42x
recall@100.98520.9786~equal
client p50447 µs595 µs1.33x
client p99644 µs1.19 ms1.85x
client p99.9798 µs1.55 ms1.95x
server p50353 µs441 µs1.25x
server p99543 µs983 µs1.81x
wall clock2.9 s4.1 s1.41x
queries sent50,00050,000
cpu18.9 s24.0 s1.27x
cpu, % of wall659%592%
waiting for a core2.2 ms14,521.1 ms6,656.05x
migrations0111,342
faults min/maj0 / 0524 / 0
peak RSS1.4 GiB3.1 GiB2.15x
disk written00
disk read00

W10-ef128 recall control, ef=128 (latency only) 1.27x

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second10,4878,2681.27x
recall@100.99400.9906~equal
client p50757 µs885 µs1.17x
client p991.14 ms1.84 ms1.62x
client p99.91.29 ms2.35 ms1.82x
server p50660 µs713 µs1.08x
server p991.04 ms1.57 ms1.52x
wall clock4.8 s6.1 s1.27x
queries sent50,00050,000
cpu33.3 s36.6 s1.10x
cpu, % of wall693%602%
waiting for a core3.8 ms26,270.7 ms6,914.82x
migrations0117,455
faults min/maj0 / 0214 / 0
peak RSS1.4 GiB3.1 GiB2.15x
disk written00
disk read00

W10-ef256 recall control, ef=256 (latency only) 1.12x

ef=256 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second6,0125,3801.12x
recall@100.99760.9956~equal
client p501.32 ms1.36 ms1.03x
client p992.02 ms2.78 ms1.38x
client p99.92.28 ms3.48 ms1.53x
server p501.22 ms1.16 ms0.95x
server p991.91 ms2.42 ms1.27x
wall clock8.4 s9.3 s1.12x
queries sent50,00050,000
cpu58.8 s58.6 s~equal
cpu, % of wall704%628%
waiting for a core4.7 ms38,470.1 ms8,170.80x
migrations0143,422
faults min/maj0 / 0179 / 0
peak RSS1.4 GiB3.1 GiB2.15x
disk written00
disk read00

W10-ef512 recall control, ef=512 (latency only) parity

ef=512 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second3,4233,373parity
recall@100.99890.9979~equal
client p502.33 ms2.24 ms0.96x
client p993.56 ms4.18 ms1.17x
client p99.94.06 ms5.27 ms1.30x
server p502.23 ms2.00 ms0.90x
server p993.45 ms3.80 ms1.10x
wall clock14.6 s14.9 s~equal
queries sent50,00050,000
cpu103.6 s96.4 s0.93x
cpu, % of wall708%649%
waiting for a core8.4 ms53,106.6 ms6,359.74x
migrations0169,278
faults min/maj0 / 0293 / 0
peak RSS1.4 GiB3.1 GiB2.15x
disk written00
disk read00
within the ±2.0% band this dataset's noise floor puts on W10-ef512: no measured difference, not a small one

W12-upload filtered search: load 100,000 with payloads

upload · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=100,000
metricstrawmANNQdrant
wall clock12.6 s24.3 s1.94x
of which upload0.5 s1.1 s2.36x
of which index wait10.0 s21.1 s2.10x
time to Green10.0 s21.1 s2.10x
cpu56.5 s85.5 s1.51x
cpu, % of wall450%351%
waiting for a core24.9 ms1,653.3 ms66.34x
migrations03,459
faults min/maj63,572 / 0112,773 / 154,555
peak RSS1.5 GiB3.1 GiB2.03x
disk written195.4 MiB837.4 MiB
disk read022.1 MiB

W12-sel1 filtered search, one keyword (~1% of bench12)

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second7,0963,368-
recall@101.00000.9999
client p50271 µs591 µs
client p99449 µs676 µs
client p99.9698 µs751 µs
server p50184 µs486 µs
server p99335 µs555 µs
wall clock7.1 s14.9 s
queries sent50,00050,000
cpu11.2 s29.7 s
cpu, % of wall159%200%
waiting for a core10.8 ms15.7 ms
migrations09,230
faults min/maj1 / 0621 / 0
peak RSS1.5 GiB3.1 GiB
disk written00
disk read00
strawmANN pass 1 -16% against two passes that agree: one pass set the spread on this row, so a band built from it is not a noise band filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set

W12-sel10 filtered search, any of 10 keywords (~10% of bench12)

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,3221,803-
recall@101.00000.9854
client p501.51 ms1.11 ms
client p991.64 ms1.34 ms
client p99.91.70 ms1.42 ms
server p501.41 ms992 µs
server p991.53 ms1.22 ms
wall clock37.9 s27.8 s
queries sent50,00050,000
cpu71.2 s56.2 s
cpu, % of wall188%202%
waiting for a core7.6 ms13.6 ms
migrations017,354
faults min/maj0 / 0381 / 0
peak RSS1.5 GiB3.1 GiB
disk written00
disk read00
recall unequal: 1.0000 vs 0.9854; §7.4 compares at equal recall filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set

W12-sel1-ef32 filtered recall control, one keyword, ef=32 (latency only) 1.26x

ef=32 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second7,2915,7741.26x
recall@101.00000.9934~equal
client p50269 µs341 µs1.27x
client p99339 µs413 µs1.22x
client p99.9560 µs489 µs0.87x
server p50183 µs239 µs1.30x
server p99241 µs289 µs1.20x
wall clock6.9 s8.7 s1.26x
queries sent50,00050,000
cpu11.0 s17.2 s1.56x
cpu, % of wall160%198%
waiting for a core1.2 ms12.5 ms10.26x
migrations09,810
faults min/maj0 / 036 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00

W12-sel1-ef64 filtered recall control, one keyword, ef=64 (latency only) 1.62x

ef=64 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second7,3124,5271.62x
recall@101.00000.9988~equal
client p50269 µs438 µs1.63x
client p99335 µs510 µs1.52x
client p99.9569 µs559 µs~equal
server p50184 µs336 µs1.82x
server p99242 µs391 µs1.61x
wall clock6.9 s11.1 s1.61x
queries sent50,00050,000
cpu11.0 s22.2 s2.01x
cpu, % of wall160%200%
waiting for a core2.1 ms13.0 ms6.31x
migrations09,451
faults min/maj0 / 062 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00

W12-sel1-ef128 filtered recall control, one keyword, ef=128 (latency only) 2.17x

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second7,2553,3392.17x
recall@101.00000.9999~equal
client p50271 µs596 µs2.19x
client p99331 µs689 µs2.08x
client p99.9504 µs761 µs1.51x
server p50185 µs491 µs2.65x
server p99235 µs564 µs2.40x
wall clock6.9 s15.0 s2.17x
queries sent50,00050,000
cpu11.1 s30.2 s2.73x
cpu, % of wall159%201%
waiting for a core1.9 ms16.8 ms8.85x
migrations010,052
faults min/maj0 / 0251 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00
2.17x is 2.72x less work per query x 0.80x cores busy during the row (1.60 against 2.00) x 1.00x clock: the middle term is occupancy, not search speed

W12-sel1-ef256 filtered recall control, one keyword, ef=256 (latency only)

ef=256 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second7,0762,441-
recall@101.00001.0000
client p50273 µs815 µs
client p99372 µs928 µs
client p99.9629 µs998 µs
server p50186 µs709 µs
server p99270 µs801 µs
wall clock7.1 s20.5 s
queries sent50,00050,000
cpu11.3 s41.3 s
cpu, % of wall159%201%
waiting for a core2.2 ms22.1 ms
migrations010,027
faults min/maj0 / 0218 / 0
peak RSS1.5 GiB3.1 GiB
disk written00
disk read00
strawmANN pass 1 -39% against two passes that agree: one pass set the spread on this row, so a band built from it is not a noise band

W12-sel1-ef512 filtered recall control, one keyword, ef=512 (latency only) 4.04x

ef=512 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second6,8331,6934.04x
recall@101.00001.0000~equal
client p50284 µs1.18 ms4.14x
client p99402 µs1.33 ms3.29x
client p99.9629 µs1.46 ms2.32x
server p50195 µs1.07 ms5.47x
server p99283 µs1.19 ms4.22x
wall clock7.4 s29.6 s4.02x
queries sent50,00050,000
cpu11.7 s59.5 s5.07x
cpu, % of wall159%201%
waiting for a core1.8 ms41.6 ms22.87x
migrations011,588
faults min/maj0 / 0313 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00
4.04x is 5.32x less work per query x 0.79x cores busy during the row (1.59 against 2.00) x 1.00x clock x 0.96x counted on-CPU share: the middle term is occupancy, not search speed

W12-sel10-ef32 filtered recall control, any of 10 keywords, ef=32 (latency only)

ef=32 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second3,2123,980-
recall@100.98080.8164
client p50627 µs495 µs
client p99775 µs638 µs
client p99.9935 µs716 µs
server p50534 µs387 µs
server p99677 µs522 µs
wall clock15.6 s12.6 s
queries sent50,00050,000
cpu27.7 s25.5 s
cpu, % of wall178%202%
waiting for a core3.0 ms10.8 ms
migrations012,232
faults min/maj0 / 0195 / 0
peak RSS1.5 GiB3.1 GiB
disk written00
disk read00
recall unequal: 0.9808 vs 0.8164; §7.4 compares at equal recall

W12-sel10-ef64 filtered recall control, any of 10 keywords, ef=64 (latency only)

ef=64 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,0772,780-
recall@100.99320.9326
client p50974 µs714 µs
client p991.17 ms886 µs
client p99.91.23 ms979 µs
server p50880 µs603 µs
server p991.07 ms765 µs
wall clock24.1 s18.0 s
queries sent50,00050,000
cpu44.7 s36.5 s
cpu, % of wall185%202%
waiting for a core2.1 ms13.1 ms
migrations014,006
faults min/maj0 / 0292 / 0
peak RSS1.5 GiB3.1 GiB
disk written00
disk read00
recall unequal: 0.9932 vs 0.9326; §7.4 compares at equal recall

W12-sel10-ef128 filtered recall control, any of 10 keywords, ef=128 (latency only)

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,3191,800-
recall@101.00000.9854
client p501.52 ms1.11 ms
client p991.65 ms1.35 ms
client p99.91.73 ms1.45 ms
server p501.41 ms994 µs
server p991.53 ms1.23 ms
wall clock37.9 s27.8 s
queries sent50,00050,000
cpu71.3 s56.2 s
cpu, % of wall188%202%
waiting for a core8.4 ms16.4 ms
migrations017,649
faults min/maj0 / 0260 / 0
peak RSS1.5 GiB3.1 GiB
disk written00
disk read00
recall unequal: 1.0000 vs 0.9854; §7.4 compares at equal recall

W12-sel10-ef256 filtered recall control, any of 10 keywords, ef=256 (latency only) 1.29x

ef=256 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,3191,0241.29x
recall@101.00000.9972~equal
client p501.52 ms1.95 ms1.29x
client p991.64 ms2.37 ms1.44x
client p99.91.73 ms2.50 ms1.44x
server p501.41 ms1.82 ms1.29x
server p991.53 ms2.23 ms1.46x
wall clock37.9 s48.9 s1.29x
queries sent50,00050,000
cpu71.3 s98.6 s1.38x
cpu, % of wall188%202%
waiting for a core5.8 ms53.8 ms9.24x
migrations019,972
faults min/maj0 / 0756 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00

W12-sel10-ef512 filtered recall control, any of 10 keywords, ef=512 (latency only) 2.16x

ef=512 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,3206122.16x
recall@101.00000.9993~equal
client p501.51 ms3.26 ms2.15x
client p991.65 ms3.90 ms2.36x
client p99.91.72 ms4.12 ms2.39x
server p501.41 ms3.08 ms2.18x
server p991.53 ms3.68 ms2.41x
wall clock37.9 s81.8 s2.16x
queries sent50,00050,000
cpu71.2 s162.6 s2.28x
cpu, % of wall188%199%
waiting for a core4.7 ms91.6 ms19.56x
migrations021,393
faults min/maj0 / 01,381 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00

W13 scroll / pagination 1.09x

closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=200,000
metricstrawmANNQdrant
queries/second50,28246,2731.09x
client p50139 µs159 µs1.15x
client p99235 µs296 µs1.26x
client p99.9275 µs440 µs1.60x
server p5011 µs28 µs2.58x
server p9916 µs93 µs5.82x
wall clock4.1 s2.2 s0.54x
queries sent200,000100,000
cpu5.6 s9.1 s1.62x
cpu, % of wall137%409%
waiting for a core2.5 ms922.2 ms369.09x
migrations0113,126
faults min/maj0 / 0622 / 0
peak RSS1.5 GiB3.1 GiB2.03x
disk written00
disk read00
strawmANN: -n 200000 (4x the table's 50000) Qdrant: -n 100000 (2x the table's 50000) strawmANN: 0.13 ms of the client's 0.14 ms p50 is not server time (13x), so 92% of this row is the load generator and the socket Qdrant: 0.13 ms of the client's 0.16 ms p50 is not server time (6x), so 82% of this row is the load generator and the socket 1.09x is 3.29x less work per query x 0.33x cores busy during the row (1.38 against 4.24) x 1.00x clock x 1.01x counted on-CPU share: the middle term is occupancy, not search speed

W11-steady mixed read/write below the rebuild threshold: search bench2 while 5,000 synthetic points append

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=330,611
metricstrawmANNQdrant
queries/second10,1027,170-
client p50771 µs932 µs
client p991.34 ms2.26 ms
client p99.93.00 ms3.83 ms
server p50674 µs744 µs
server p991.14 ms1.87 ms
wall clock33.3 s44.4 s
queries sent330,611330,611
cpu228.5 s277.1 s
cpu, % of wall685%624%
waiting for a core3,877.6 ms247,440.9 ms
migrations16894,806
faults min/maj23,682 / 0376,849 / 480,766
peak RSS1.5 GiB3.3 GiB
disk written9.9 MiB3.7 GiB
disk read095.2 MiB
strawmANN: qps over the 25.0 s of the append the search saw (10,152 over the whole search); append 200 points/s Qdrant: qps over the 25.0 s of the append the search saw (7,523 over the whole search); append 200 points/s search-during-write; no recall join

W11 mixed read/write: search bench2 while 20,000 synthetic points append (runs last)

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=863,509
metricstrawmANNQdrant
queries/second9,5267,151-
client p50809 µs937 µs
client p991.98 ms2.25 ms
client p99.93.91 ms6.85 ms
server p50711 µs753 µs
server p991.43 ms1.87 ms
wall clock92.1 s116.7 s
queries sent863,509863,509
cpu634.6 s720.3 s
cpu, % of wall689%617%
waiting for a core18,357.3 ms572,136.9 ms
migrations602,221,296
faults min/maj89,971 / 0658,653 / 439,853
peak RSS1.5 GiB5.0 GiB
disk written39.4 MiB5.8 GiB
disk read0147.3 MiB
strawmANN: qps over the 66.7 s of the append the search saw (9,566 over the whole search); append 300 points/s Qdrant: qps over the 66.7 s of the append the search saw (7,528 over the whole search); append 300 points/s search-during-write; no recall join
Host discipline the §7.1 gate, per check, and ambient load per row

Ambient load per row

The §7.1 gate checks the machine once, at the start. Colour is the engine, as everywhere else; a hatched bar with a red edge is a row that had another process on the box while it ran, which the gate cannot see.

strawmANN gate pass

environment hash 373e8531fd3d1bf6
okAMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical)
okgovernor=performance
okboost disabled
oksmt=on
declared rather than required (§7.1 asks for SMT on or off, not
for off), and hashed, so these rows never share a chart with
SMT-off ones
cpu N and cpu N+12 share a core (12 pairs);
a --server-cpus or --client-cpus set holding both cpus of a
pair measures the engine on half the cores it names
okprofile=as-deployed (isolcpus=none nohz_full=none)
the scheduler is running as it ships, so these numbers describe a
deployment rather than the engine in isolation; they may not be
compared against an `isolated` run, and the environment hash
enforces that
okthp=madvise numa_nodes=1 numa_balancing=0
okperf_event_paranoid=-1
oktask_delayacct=1, so time blocked on a device is measurable
okperf at /usr/bin/perf, so `workloads.py run --perf` can attach
okcgroup io controller reaches the engines' own scope, so block-layer read/write operations are measurable
okzig=0.16.0
okmachine is quiescent (load 0.1 over 24 cores)
measurement profile: as-deployed

Qdrant gate pass

environment hash 373e8531fd3d1bf6
okAMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical)
okgovernor=performance
okboost disabled
oksmt=on
declared rather than required (§7.1 asks for SMT on or off, not
for off), and hashed, so these rows never share a chart with
SMT-off ones
cpu N and cpu N+12 share a core (12 pairs);
a --server-cpus or --client-cpus set holding both cpus of a
pair measures the engine on half the cores it names
okprofile=as-deployed (isolcpus=none nohz_full=none)
the scheduler is running as it ships, so these numbers describe a
deployment rather than the engine in isolation; they may not be
compared against an `isolated` run, and the environment hash
enforces that
okthp=madvise numa_nodes=1 numa_balancing=0
okperf_event_paranoid=-1
oktask_delayacct=1, so time blocked on a device is measurable
okperf at /usr/bin/perf, so `workloads.py run --perf` can attach
okcgroup io controller reaches the engines' own scope, so block-layer read/write operations are measurable
okzig=0.16.0
busy: Xorg(7%)
measurement profile: as-deployed
What this does not establish

No Qdrant developer has used this tool. Nothing here has been run, checked against something already known, or disagreed with by anyone outside the project: every number has one author and one reviewer, and they are the same person. A result you can refute is more useful to us than one you accept.

It has never been pointed at a real Qdrant regression. The harness detects a known ISA slowdown and correctly reports no difference on a control row. That is internal consistency, not evidence it would flag a regression in your tree or stay quiet through a refactor. bench/harness/qdrant_ab.py exists to run two of your commits through it blind, and that experiment has not been done.

Part of the measured throughput gap is kernel width, not architecture. Read from a dev checkout: Qdrant's distance path tops out at four 256-bit accumulators and has no AVX-512, while these kernels use 512-bit ones on a machine that has them. That is a real difference and it is not the same claim as "the design is faster".

The engines are not equivalent, by construction. §1 removes sharding, replication, consensus, snapshots, sparse vectors, multivectors and disk-resident operation — permanently. Payload storage and filtering were in that list until 2026-09-03 and are now built (M7), so W12 measures a filtered search with a keyword index on both engines. Anything still unbuilt answers UNIMPLEMENTED naming the construct and reports n/a, never a silent degradation. A strawman that was not faster would mean it was badly built; the question is by how much, and where the model was wrong.

Reproducing this

Everything below runs from a clone. The engine has no dependencies; the analysis path is a uv project pinned by bench/uv.lock.

scripts/doctor.py                    # what can this machine measure?
bench/setup.py check                 # §7.1 host gate: governor, boost, isolation, idle
conformance/datasets/datasets.py fetch

# one engine at a time, §7.1
zig build -Doptimize=ReleaseFast
# --connections matters: W4 opens ~32 sockets (-t 16 -c 2) and a server with
# fewer closes the excess before the HTTP/2 preface, which the client reports
# only as "transport error".
./zig-out/bin/strawmann --port 6334 --connections 64
bench/harness/workloads.py run http://localhost:6334 strawmann --sink --report

# the other half of every number: recall, and the licence to compare at all
cd conformance
cargo run --release -- relevance --engine http://localhost:6334 --label strawmann \
  --base $DATA/laion-small-clip/base.fbin --queries $DATA/laion-small-clip/queries.fbin \
  --ground-truth $DATA/laion-small-clip/gt/laion-small-clip.cosine.k100.gt.json --metric cosine \
  --ef 32 64 128 256 512 --json ../bench/results/strawmann/recall.json
cargo run --release -- differ --strawmann http://localhost:6334 --qdrant http://localhost:6434 \
  --base $DATA/laion-small-clip/base.fbin --queries $DATA/laion-small-clip/queries.fbin \
  --ground-truth $DATA/laion-small-clip/gt/laion-small-clip.cosine.k100.gt.json \
  --json ../bench/results/strawmann/conformance.json

uv run --project bench bench/harness/report.py strawmann qdrant

Disagreements are the point. docs/spec.md is what all of this is measured against, docs/bugs.md lists the measurement bugs found so far (each produced a plausible wrong number rather than a failure), and docs/validation.md records what has and has not been checked.