strawmANN benchmark report

strawmANN against Qdrant · AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical) · generated 2026-09-30 06:41:52 CEST

Dataset
dbpedia-openai-1m
Vectors
1,000,000 × 1,536 · 10,000 held-out queries
Metric
cosine
Storage datatype
strawmANN float32 · Qdrant float32 (default)
Quantization
none (fp32) · separate rows measure binary (1 bit/dim), PQ (product), SQ8 scalar
strawmANN: 43/43 ok Qdrant: 43/43 ok

Summary

At equal recall, strawmANN serves 1.09x to 1.40x Qdrant's throughput. That range spans the recall levels measured; allowing for the uncertainty in each recall figure widens it to 0.90x to 1.53x.

throughput at equal recall
1.09 to 1.40x
strawmANN over Qdrant, recall@10 0.895 to 0.989
p99 latency at 90% load
strawmANN2.83 ms
Qdrant9.53 ms
W4-sat90, open loop, one offered rate
upload and index build
strawmANN260 s
Qdrant404 s
W1 + W2
peak memory (RSS)
strawmANN31.6 GiB
Qdrant43.6 GiB
largest before the concurrent-write rows (W11); of it, anonymous 4.0 GiB / 3.4 GiB
written to disk
strawmANN29.5 GiB
Qdrant134.1 GiB
before the concurrent-write rows (W11); they wrote 1.4 GiB / 21.4 GiB more

Where strawmANN is slower, past the noise floor:

Throughput against recall

Read it vertically: at any recall both engines reach, the higher curve is faster. Throughput from bfb W10, recall from the conformance sweep, joined on ef. ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.

Every figure is the median of 3 passes per engine, run alternately (A/B/A/B). What this does not establish.

Glossary
recall@10
Of the ten nearest neighbours a query really has, the share the engine returned. 1.0 is a perfect answer. "Really has" is settled by an exhaustive fp64 search, not by the other engine.
ef
The candidate-list size: how many nodes the index keeps in play while searching. Larger is slower and more accurate. It does not mean the same amount of work in both engines, which is why the headline holds recall fixed instead.
MRDE
Mean relative distance error: when the engine returns a wrong neighbour, how much further away it is than the right one. Small numbers mean the misses were near-misses.
qps
Queries per second. On a batched row one request carries several queries, and the request rate is shown under it.
p50 / p99 / p99.9
The latency half of requests beat, that 99% beat, that 99.9% beat. The tail is what a user notices.
closed / open loop
A closed loop sends the next request only when the last one comes back, so a slow server receives less work and its tail looks better than it is. An open loop sends at a fixed rate regardless.
noise floor
How much a number moves between identical runs on this machine. A ratio inside it is shown grey with ≈: no measured difference. Hover a ratio to see its band.
conformance tier
What the two engines were shown to agree on before any speed was quoted. T1 licenses a single engine's own numbers; T3 licenses comparing the two, and requires their recall to be statistically indistinguishable.
§ numbers
Sections of docs/spec.md, the written rule each claim in the appendix is measured against. §7.1 is the host gate, §7.4 the comparison rules, §8 what may be published.

Throughput

Queries per second; higher is better. The ratio is strawmANN over Qdrant: green is faster, red slower, grey ≈ inside the noise floor (hover a ratio for its band). A dash means the pair is not compared, and the note says why. An ef sweep is one row showing its range; every measurement is under All throughput rows.

workloadstrawmANNQdrantrationotes
W3search, fp32, single query1,0606241.70x
W4search, saturating (closed loop)3,5563,4831.02x
W5search batched (16 distinct dataset queries per request)3,530221 requests/s1,30382 requests/s2.71x
W6quantized: scalar3,0282,0261.50x
W6 ef 32 to 512SQ8 recall control (latency only)1,063 to 6,051674 to 4,3611.37 to 1.58x
W7quantized: binary + oversampling4,1683,5791.16x
W8quantized: PQ1,4821,1061.34x
W9exact / brute force13101.33x
W10 ef 32 to 512recall control (latency only)1,235 to 10,6251,068 to 7,3581.16 to 1.44x
W12-sel1filtered search, one keyword (~1% of bench12)1,8691,6161.16x
W12-sel10filtered search, any of 10 keywords (~10% of bench12)822823-recall differs: 0.9893 vs 0.9720
W12-sel1 ef 32 to 512filtered recall control, one keyword (latency only)1,861 to 1,864758 to 3,3970.77 to 2.45x1 of 5 points not compared recall differs
W12-sel10 ef 32 to 512filtered recall control, any of 10 keywords (latency only)294 to 1,974286 to 2,259parity3 of 5 points not compared recall differs
W13scroll / pagination51,14146,8521.09xstrawmANN: 92% is client and socket, not server Qdrant: 82% is client and socket, not server
W11-steadymixed read/write below the rebuild threshold: search bench2 while 49,500 synthetic points append580504237 requests/s-search-during-write; no recall join
W11mixed read/write: search bench2 while 198,000 synthetic points append (runs last)25199-search-during-write; no recall join

Throughput by workload

Higher is better. Hatched bars are rows the table does not compare (W11): each bar is that engine's own rate, and the pair is not a result.

Recall

recall@10 against an exact fp64 search, not against the other engine. Equal ef is not equal work in the two engines, so the comparison that counts holds recall fixed and compares throughput there.

At matched recall

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.8945Qdrant at ef=32, strawmANN interpolated10,2737,3581.40x1.28 – 1.53x
0.9399strawmANN at ef=64, Qdrant interpolated6,4815,0371.29x1.24 – 1.34x
0.9441Qdrant at ef=64, strawmANN interpolated5,9674,8651.23x1.09 – 1.38x
0.9667strawmANN at ef=128, Qdrant interpolated3,7993,2271.18x1.10 – 1.26x
0.9694Qdrant at ef=128, strawmANN interpolated3,4273,0741.11x0.96 – 1.30x
0.9812strawmANN at ef=256, Qdrant interpolated2,1811,9331.13x1.01 – 1.26x
0.9822Qdrant at ef=256, strawmANN interpolated2,0241,8541.09x0.90 – 1.32x
0.9892strawmANN at ef=512, Qdrant interpolated1,2351,1371.09x0.93 – 1.27x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W10 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column is why the ratio is not a result on its own. The ratio treats the anchor's recall as exact; it is an estimate, and moving it across its 95% interval moves the interpolated rate with it — near the top of the sweep 0.003 of recall spans a factor of 1.8, so two decimals there quote the interpolation. Even that is the narrow reading: it moves one recall and not the other, adjacent anchors share bracketing segments, and the throughputs behind it are medians of the run's passes.

At matched recall, SQ8

The same reading over the scalar-quantized collection. Under the `pool` policy both engines rescore `ef` candidates, so W6 is compared row by row as well; held at equal recall, the reading does not depend on the rows landing in one recall band. Note where each engine's curve stops, because a recall only one of them reaches is the more useful fact about an encoding than any ratio.

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.8918strawmANN at ef=32, Qdrant interpolated6,0514,3261.40x1.34 – 1.46x
0.9393strawmANN at ef=64, Qdrant interpolated4,2663,1501.35x1.31 – 1.40x
0.9413Qdrant at ef=64, strawmANN interpolated4,1613,1091.34x1.24 – 1.45x
0.9669strawmANN at ef=128, Qdrant interpolated3,0222,0611.47x1.38 – 1.56x
0.9682Qdrant at ef=128, strawmANN interpolated2,8882,0181.43x1.22 – 1.68x
0.9811strawmANN at ef=256, Qdrant interpolated1,8411,2041.53x1.36 – 1.72x
0.9820Qdrant at ef=256, strawmANN interpolated1,7391,1641.49x1.19 – 1.88x
0.9895strawmANN at ef=512, Qdrant interpolated1,0637001.52x1.31 – 1.76x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W6 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column reads as it does in the fp32 table above.

At matched recall, filtered to 10%

The same reading under a keyword filter over the points the condition matched, 19,997 to 20,119 across 3 builds in strawmANN and 19,982 to 20,188 across 3 builds in Qdrant (bfb draws the keyword payloads unseeded at each upload). The per-row table refuses W12-sel10 a ratio because the two engines land just outside the recall band at the one ef it measures; held at equal recall instead, the comparison exists at every recall both engines reach. Note the shape rather than any single number: one engine's curve is flat in ef and the other's is steep, so where you match decides the ratio.

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.9383strawmANN at ef=32, Qdrant interpolated1,9741,0561.87x1.80 – 1.94x
0.9720Qdrant at ef=128, strawmANN interpolated1,3258221.61x1.54 – 1.68x
0.9728strawmANN at ef=64, Qdrant interpolated1,3128071.63x1.48 – 1.79x
0.9893strawmANN at ef=128, Qdrant interpolated8225491.50x1.41 – 1.59x
0.9945Qdrant at ef=256, strawmANN interpolated5364861.10x0.96 – 1.27x
0.9959strawmANN at ef=256, Qdrant interpolated4804151.16x0.98 – 1.37x
0.9991Qdrant at ef=512, strawmANN interpolated3282861.15x1.06 – 1.24x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W12-sel10 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column reads as it does in the fp32 table above.

Recall by ef

efstrawmANNQdrant
recall@1recall@10recall@100MRDErecall@1recall@10recall@100MRDE
ef=320.89490.8912-1.86e-030.90060.8945-1.78e-03
ef=640.93780.9399-8.97e-040.93980.9441-8.27e-04
ef=1280.96580.96670.92914.73e-040.96720.96940.93144.27e-04
ef=2560.98030.98120.96582.54e-040.98100.98220.96782.44e-04
ef=5120.98910.98920.98341.43e-040.98990.99010.98481.34e-04
How recall was measured

10,000 held-out queries, limit 10, ε=9.537e-07, against the fp64 oracle rather than against the other engine. Base checksum 71f1df815731cdbf: the same corpus the latency rows were measured on. Each pass rebuilds the collection. Across 3 independent builds of it, recall@10 at ef=512 spread by: strawmANN 0.00035; Qdrant 0.00033. Over the same builds the engine reported unreachable nodes: strawmANN 0 of 200,000 and 0 of 990,000 — a node nothing points at is invisible at any ef, so that is where a graph-quality difference shows rather than being inferred from the recall beside it. Level seed: strawmANN 0x57ea3111. Recorded as provenance: under this engine's level draw, six builds at four seeds sit within 0.00003 of recall@10 at ef 512, well inside the spread above.

What each quantization costs in recall

Colour is the encoding here, not the engine — the engines are the line style. Each sweep is joined on the quantization parameters it actually sent, so a curve speaks only for the search its row ran. Higher is better, and the encodings are not free: read this against the throughput their rows bought.

Latency

Client-side round trip. A closed loop understates the tail (a stalled server stops receiving requests), so this keeps the fixed-rate, open-loop rows, which offer the same load to both engines, plus single-client W3 and batched W5. Every row is under All latency rows.

workloadstrawmANNQdrant
p50p99p99.9p50p99p99.9
W3 search, fp32, single query947 µs1.37 ms1.89 ms1.60 ms2.29 ms2.57 ms
W4-sat50 search, fixed rate at 50% of saturation (open loop)1.03 ms1.63 ms1.87 ms1.76 ms2.95 ms3.77 ms
W4-sat70 search, fixed rate at 70% of saturation (open loop)1.11 ms1.95 ms2.29 ms2.00 ms3.75 ms4.76 ms
W4-sat90 search, fixed rate at 90% of saturation (open loop)1.39 ms2.83 ms3.56 ms2.70 ms9.53 ms17.91 ms
W5 search batched (16 distinct dataset queries per request)9.06 ms10.55 ms11.31 ms24.54 ms27.79 ms28.86 ms

Ingest and index build

Qdrant indexes while it ingests, so its upload time already contains most of the indexing; strawmANN uploads raw and builds afterwards. Compare the sum, not the upload line.

Where the load time goes

Solid is upload, hatched is the wait for Green. The bar's whole length is the sum this section asks you to compare; the split is why the upload line alone is not comparable between these engines.

workloadstrawmANNQdrant
W0-uploadd=4 floor: graph traversal with the distance taken out17.05 s22.05 s
W1ingest throughput (no index wait)5.53 s33.59 s
W2index build, time to Green254.86 s370.89 s
W6-uploadsearch, SQ8 scalar quantization258.85 s176.40 s
W7-uploadsearch, binary quantization263.93 s89.19 s
W8-uploadsearch, product quantization359.29 s875.15 s
W12-uploadupload39.12 s90.22 s
W1 + W2upload and index, together260.4 s404.5 s1.55x

Memory and disk

What each engine held in memory and moved to and from disk over the whole run. Per-row figures are under Storage and I/O.

strawmANNQdrant
storage on disk35.5 GiB26.5 GiB
peak memory (RSS)31.6 GiB43.6 GiB
read from disk86.0 MiB1.2 GiB
written to disk29.5 GiB134.1 GiB

Peak memory and storage on disk are read before the concurrent-write rows (W11): an engine rewriting segments maps old and new files at once, and RSS counts each mapping. Bytes written are summed over the same rows, since W11's volume is set by the harness's write rate; the other disk figures are totals over every row.

Appendix

How the run was set up, every row of every table, and the diagnostics behind them. Charts on a logarithmic axis say so on the axis: read the positions there, not the distances.

Run conditions and conformance

Measured as-deployed: the scheduler was left as it ships, so a difference is what a user would see rather than the engine in isolation, and ambient load is part of the measurement. Qdrant ran equal-work: asked for one populated graph (--segments 1), as strawmANN serves, so ef means the same thing on both sides. This is the configuration §8's comparative licensing is built around, and it is a control rather than a deployment. Its optimizer treats the count as a target; what it actually held is read back from the engine and stated under Collection config.

Conformance tier reached: T4 quantization fidelity. T1 passed, so single-engine performance rows are licensed (§8). T3 passed, so a strawmANN-vs-Qdrant throughput comparison is licensed: the engines are at equal recall (§7.4). Conformance hash 2880404844a92b92.

On exact search the two engines' scores differ by at most 3.576e-07 (p99 1.788e-07), against a calibrated absolute ε of 9.537e-07. §8.1: bit-exactness is unachievable between two different summation orders, so this — equal values within a measured tolerance — is the claim the speed rests on.

The conformance harness's single-client rate does not order the engines consistently (0.71x to 2.73x); a different instrument, so no ratio is formed from it.

Every figure is the median of 3 passes per row per engine, alternated A/B/A/B (§7.4), and the spread is this run's own over 36 of 43 rows. The ingest and index-build rows (W0-upload, W1, W2, W6-upload, W7-upload, W8-upload, W12-upload) have none, so no build time carries a verdict. Each pass rebuilds the index, so the bands include build variance and are wider than a floor measured against one standing graph.

Units are queries per second. bfb reports rps, which counts batch requests: at --search-batch-size 16 the two differ by 16x, and reading one as the other once turned a 2.2x speedup into an apparent 7x regression. The tables show the request rate wherever it diverges.

The open-loop rows are not a speed. W4-sat50/70/90 offer a fixed fraction of measured saturation; serving it means the engine kept up, not that it was faster, so no ratio is printed. They exist because --parallel is a closed loop, where a stalled server stops receiving requests and understates its own tail.

ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.

What was measured builds, dataset, host
strawmANN
sm-dbp1m-perf-0930
Qdrant
qd-dbp1m-perf-0930
enginecommit 1fb2e437e858
ReleaseFast, native build, vnni on
binary sha256 e9f47255fbea5cee
built 2026-09-29T15:06:47Z with zig 0.16.0
version 1.19.2-dev
binary ~/.cache/strawmann/qdrant-dbeb0f73/qdrant
sha256 dbeb0f73dea2d371 build 878843e6 (from the server's banner, not a checkout)
native binary outside a checkout: sha256 is the identity (§8.9)
network native no container in the path, like strawmann
measured2026-09-29T20:30:39Z to 2026-09-30T02:27:34Z (3 passes)2026-09-29T21:34:47Z to 2026-09-30T03:49:45Z (3 passes)
profileas-deployed · both engines
§7.1 gatepass · both engines
environment hash373e8531fd3d1bf6 · both engines
load generatorbfb dev @ fc6632e5 (qdrant/bfb#176; carries #172, findings 32's --rps reaping fix) · both engines
clientqdrant-client 1.16.1-dev (git dev branch) · both engines
launched as~/Workspace/strawmann/zig-out/bin/strawmann --port 6334 --capacity 1237500 --connections 64 --workers 7 --io-threads 1 --pin --cpus 4-11 --data-dir ~/.cache/strawmann/strawmann-storage --default-placement cached~/.cache/strawmann/qdrant-dbeb0f73/qdrant
cores the engine could use4-11 (8 cores) requested by the harness
4-11 (8 cores) observed on 8 threads
0-23 (24 cores) observed on 1 thread
4-11 (8 cores) requested by the harness
4-11 (8 cores) observed on 50 threads

Dataset

dbpedia-openai-1m, 1000000 × 1536, cosine, 10000 held-out queries
ground truth: recomputed by our fp64 oracle
26 files, each pinned by sha256 in datasets.json

Host

AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M
24 logical cores · 58.6 GB · kernel 7.0.0-34-generic
memory bandwidth: 73.6 GB/s aggregate (24 threads) · 45.9 GB/s single core (62% of bus)
strawmann startup banner, 2026-09-29T20:30:40Z; the Qdrant run agrees within 10%
All throughput rows every measurement, with its notes
workloadstrawmANNQdrantrationotes
W0d=4 floor: graph traversal with the distance taken out4,5983,4951.32x1.32x is 1.83x less work per query x 0.67x cores busy during the row (0.67 against 1.01) x 1.00x clock x 1.08x counted on-CPU share: the middle term is occupancy, not search speed d=4 floor: this measures graph traversal and heaps with the distance taken out, not request plumbing (findings 28: 76% of the server's CPU was HNSW search, 24% kernel; the RPC path is a small remainder)
W3search, fp32, single query1,0606241.70x
W4search, saturating (closed loop)3,5563,4831.02x
W4-sat50search, fixed rate at 50% of saturation (open loop)1,7411,741offered
W4-sat70search, fixed rate at 70% of saturation (open loop)2,4372,437offered
W4-sat90search, fixed rate at 90% of saturation (open loop)3,1333,133offered
W5search batched (16 distinct dataset queries per request)3,530221 requests/s1,30382 requests/s2.71xper-batch latency (16 queries/request); queries: dataset, random-sample 2.71x is 0.75x less work per query x 3.58x cores busy during the row (7.07 against 1.97) x 1.00x clock: the middle term is occupancy, not search speed batched 16, so qps and rps differ by 16x; the ratio uses qps
W6quantized: scalar3,0282,0261.50x
W6-ef32SQ8 recall control, ef=32 (latency only)6,0514,3611.39x1.39x is 1.70x less work per query x 0.79x cores busy during the row (1.59 against 2.02) x 1.00x clock x 1.04x counted on-CPU share: the middle term is occupancy, not search speed
W6-ef64SQ8 recall control, ef=64 (latency only)4,2663,1091.37x
W6-ef128SQ8 recall control, ef=128 (latency only)3,0222,0181.50x
W6-ef256SQ8 recall control, ef=256 (latency only)1,8411,1641.58x
W6-ef512SQ8 recall control, ef=512 (latency only)1,0636741.58x
W7quantized: binary + oversampling4,1683,5791.16x
W8quantized: PQ1,4821,1061.34x
W9exact / brute force13101.33x1.33x is 2.62x less work per query x 0.53x cores busy during the row (4.20 against 7.98) x 1.00x clock x 0.97x counted on-CPU share: the middle term is occupancy, not search speed brute force over the whole collection, no index involved
W10-ef32recall control, ef=32 (latency only)10,6257,3581.44x
W10-ef64recall control, ef=64 (latency only)6,4814,8651.33x
W10-ef128recall control, ef=128 (latency only)3,7993,0741.24x
W10-ef256recall control, ef=256 (latency only)2,1811,8541.18x
W10-ef512recall control, ef=512 (latency only)1,2351,0681.16x
W12-sel1filtered search, one keyword (~1% of bench12)1,8691,6161.16xfiltered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set
W12-sel10filtered search, any of 10 keywords (~10% of bench12)822823-recall unequal: 0.9893 vs 0.9720; §7.4 compares at equal recall filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set
W12-sel1-ef32filtered recall control, one keyword, ef=32 (latency only)1,8643,397-recall unequal: 1.0000 vs 0.9804; §7.4 compares at equal recall
W12-sel1-ef64filtered recall control, one keyword, ef=64 (latency only)1,8622,4120.77x
W12-sel1-ef128filtered recall control, one keyword, ef=128 (latency only)1,8641,6361.14x
W12-sel1-ef256filtered recall control, one keyword, ef=256 (latency only)1,8641,0831.72x
W12-sel1-ef512filtered recall control, one keyword, ef=512 (latency only)1,8617582.45x
W12-sel10-ef32filtered recall control, any of 10 keywords, ef=32 (latency only)1,9742,259-recall unequal: 0.9383 vs 0.7668; §7.4 compares at equal recall
W12-sel10-ef64filtered recall control, any of 10 keywords, ef=64 (latency only)1,3121,400-recall unequal: 0.9728 vs 0.9005; §7.4 compares at equal recall
W12-sel10-ef128filtered recall control, any of 10 keywords, ef=128 (latency only)822822-recall unequal: 0.9893 vs 0.9720; §7.4 compares at equal recall
W12-sel10-ef256filtered recall control, any of 10 keywords, ef=256 (latency only)480486paritywithin the ±4.5% band this dataset's noise floor puts on W12-sel10-ef256: no measured difference, not a small one
W12-sel10-ef512filtered recall control, any of 10 keywords, ef=512 (latency only)294286paritywithin the ±4.9% band this dataset's noise floor puts on W12-sel10-ef512: no measured difference, not a small one
W13scroll / pagination51,14146,8521.09xstrawmANN: -n 200000 (4x the table's 50000) Qdrant: -n 100000 (2x the table's 50000) strawmANN: 0.13 ms of the client's 0.14 ms p50 is not server time (12x), so 92% of this row is the load generator and the socket Qdrant: 0.13 ms of the client's 0.16 ms p50 is not server time (6x), so 82% of this row is the load generator and the socket 1.09x is 3.25x less work per query x 0.33x cores busy during the row (1.40 against 4.23) x 1.00x clock x 1.02x counted on-CPU share: the middle term is occupancy, not search speed
W11-steadymixed read/write below the rebuild threshold: search bench2 while 49,500 synthetic points append580504237 requests/s-strawmANN: qps over the 26.1 s the append ran (726 over the whole search); append 1,900 points/s Qdrant: qps over the 26.0 s the append ran (272 over the whole search); append 1,900 points/s search-during-write; no recall join
W11mixed read/write: search bench2 while 198,000 synthetic points append (runs last)25199-strawmANN: qps over the 31.7 s the append ran (251 over the whole search); write overlap 53%; search covered 53% of the append; append 3,300 points/s Qdrant: qps over the 82.1 s the append ran (97 over the whole search); search covered 97% of the append; append 2,396 points/s search-during-write; no recall join; the search saw under 90% of strawmANN's append, so its qps is the append's start and not all of it

W10: throughput against ef

ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.
All latency rows closed loop included, p50 to max

A row flagged “… not server time” is one where the client saw far more than the server reported: below the rate at which a queue can form, what is left is the load generator's own scheduling.

workloadstrawmANN p50strawmANN p95strawmANN p99strawmANN p99.9strawmANN maxQdrant p50Qdrant p95Qdrant p99Qdrant p99.9Qdrant max
W0d=4 floor: graph traversal with the distance taken out closed loop 213 µs240 µs261 µs305 µs891 µs285 µs319 µs348 µs400 µs1.44 ms
W3search, fp32, single query closed loop 947 µs1.21 ms1.37 ms1.89 ms2.58 ms1.60 ms2.14 ms2.29 ms2.57 ms4.63 ms
W4search, saturating (closed loop) closed loop 18.05 ms19.55 ms20.22 ms26.83 ms39.91 ms18.14 ms23.96 ms32.49 ms43.64 ms67.50 ms
W4-sat50search, fixed rate at 50% of saturation (open loop) open loop 1.03 ms1.38 ms1.63 ms1.87 ms5.92 ms1.76 ms2.41 ms2.95 ms3.77 ms8.13 ms
W4-sat70search, fixed rate at 70% of saturation (open loop) open loop 1.11 ms1.57 ms1.95 ms2.29 ms6.69 ms2.00 ms3.11 ms3.75 ms4.76 ms14.11 ms
W4-sat90search, fixed rate at 90% of saturation (open loop) open loop 1.39 ms2.36 ms2.83 ms3.56 ms7.69 ms2.70 ms4.89 ms9.53 ms17.91 ms35.03 ms
W5search batched (16 distinct dataset queries per request) closed loop 9.06 ms10.13 ms10.55 ms11.31 ms15.90 ms24.54 ms26.76 ms27.79 ms28.86 ms29.30 ms
W6quantized: scalar closed loop 663 µs802 µs867 µs1.00 ms2.07 ms987 µs1.23 ms1.30 ms1.43 ms2.93 ms
W6-ef32SQ8 recall control, ef=32 (latency only) closed loop 313 µs497 µs587 µs705 µs1.27 ms454 µs552 µs617 µs715 µs1.76 ms
W6-ef64SQ8 recall control, ef=64 (latency only) closed loop 458 µs623 µs826 µs981 µs1.52 ms642 µs776 µs842 µs998 µs2.62 ms
W6-ef128SQ8 recall control, ef=128 (latency only) closed loop 665 µs803 µs866 µs984 µs2.07 ms991 µs1.23 ms1.30 ms1.43 ms2.93 ms
W6-ef256SQ8 recall control, ef=256 (latency only) closed loop 1.09 ms1.34 ms1.41 ms1.56 ms3.11 ms1.71 ms2.20 ms2.31 ms2.45 ms4.43 ms
W6-ef512SQ8 recall control, ef=512 (latency only) closed loop 1.88 ms2.36 ms2.48 ms2.63 ms4.96 ms2.95 ms3.86 ms4.06 ms4.25 ms6.84 ms
W7quantized: binary + oversampling closed loop 459 µs611 µs670 µs767 µs1.28 ms558 µs628 µs671 µs777 µs1.91 ms
W8quantized: PQ closed loop 1.36 ms1.64 ms1.77 ms1.99 ms2.49 ms1.81 ms2.15 ms2.32 ms2.60 ms3.94 ms
W9exact / brute force closed loop 575.52 ms1159.76 ms1504.90 ms2481.43 ms3291.95 ms808.83 ms856.44 ms871.60 ms886.78 ms890.70 ms
W10-ef32recall control, ef=32 (latency only) closed loop 737 µs1.02 ms1.20 ms1.67 ms5.82 ms1.00 ms1.70 ms1.98 ms2.54 ms6.33 ms
W10-ef64recall control, ef=64 (latency only) closed loop 1.23 ms1.67 ms1.92 ms2.39 ms7.25 ms1.54 ms2.54 ms3.02 ms3.73 ms7.85 ms
W10-ef128recall control, ef=128 (latency only) closed loop 2.11 ms2.89 ms3.27 ms3.76 ms9.74 ms2.52 ms3.93 ms4.63 ms5.69 ms8.31 ms
W10-ef256recall control, ef=256 (latency only) closed loop 3.66 ms5.14 ms5.78 ms6.60 ms11.65 ms4.23 ms6.30 ms7.42 ms8.94 ms12.10 ms
W10-ef512recall control, ef=512 (latency only) closed loop 6.44 ms9.18 ms10.35 ms11.85 ms18.28 ms7.39 ms10.78 ms12.24 ms14.06 ms17.58 ms
W12-sel1filtered search, one keyword (~1% of bench12) closed loop 1.06 ms1.11 ms1.13 ms1.19 ms2.89 ms1.24 ms1.38 ms1.44 ms1.51 ms3.52 ms
W12-sel10filtered search, any of 10 keywords (~10% of bench12) closed loop 2.49 ms2.89 ms3.03 ms3.21 ms4.27 ms2.44 ms2.68 ms2.78 ms2.94 ms5.10 ms
W12-sel1-ef32filtered recall control, one keyword, ef=32 (latency only) closed loop 1.07 ms1.11 ms1.13 ms1.18 ms2.97 ms585 µs651 µs687 µs752 µs2.05 ms
W12-sel1-ef64filtered recall control, one keyword, ef=64 (latency only) closed loop 1.07 ms1.11 ms1.13 ms1.19 ms2.81 ms826 µs917 µs956 µs1.02 ms2.68 ms
W12-sel1-ef128filtered recall control, one keyword, ef=128 (latency only) closed loop 1.07 ms1.11 ms1.13 ms1.17 ms2.84 ms1.22 ms1.36 ms1.41 ms1.49 ms3.41 ms
W12-sel1-ef256filtered recall control, one keyword, ef=256 (latency only) closed loop 1.07 ms1.11 ms1.13 ms1.17 ms2.85 ms1.85 ms2.02 ms2.09 ms2.19 ms4.36 ms
W12-sel1-ef512filtered recall control, one keyword, ef=512 (latency only) closed loop 1.07 ms1.11 ms1.13 ms1.17 ms2.92 ms2.63 ms2.81 ms2.88 ms3.17 ms5.45 ms
W12-sel10-ef32filtered recall control, any of 10 keywords, ef=32 (latency only) closed loop 1.02 ms1.19 ms1.32 ms1.56 ms2.31 ms872 µs1.04 ms1.14 ms1.30 ms2.80 ms
W12-sel10-ef64filtered recall control, any of 10 keywords, ef=64 (latency only) closed loop 1.55 ms1.79 ms1.90 ms2.16 ms3.02 ms1.42 ms1.62 ms1.73 ms1.89 ms3.71 ms
W12-sel10-ef128filtered recall control, any of 10 keywords, ef=128 (latency only) closed loop 2.49 ms2.89 ms3.03 ms3.22 ms4.27 ms2.44 ms2.68 ms2.79 ms2.93 ms5.13 ms
W12-sel10-ef256filtered recall control, any of 10 keywords, ef=256 (latency only) closed loop 4.23 ms4.98 ms5.27 ms5.55 ms6.05 ms4.13 ms4.57 ms4.75 ms5.11 ms8.01 ms
W12-sel10-ef512filtered recall control, any of 10 keywords, ef=512 (latency only) closed loop 6.81 ms7.30 ms7.56 ms7.85 ms11.51 ms6.99 ms7.83 ms8.12 ms8.46 ms10.79 ms
W13scroll / pagination closed loop 137 µs203 µs230 µs267 µs1.11 ms157 µs219 µs293 µs446 µs1.44 ms
W11-steadymixed read/write below the rebuild threshold: search bench2 while 49,500 synthetic points append closed loop 9.02 ms26.38 ms34.05 ms114.14 ms158.21 ms31.01 ms56.46 ms68.87 ms82.97 ms107.60 ms
W11mixed read/write: search bench2 while 198,000 synthetic points append (runs last) closed loop 21.89 ms91.16 ms113.84 ms140.08 ms163.20 ms70.85 ms152.35 ms186.56 ms1871.42 ms1927.46 ms
Storage and I/O block layer and syscalls, per row

The syscall rows count every descriptor, sockets included, so on a search row they measure the network rather than the disk.

strawmANNQdrant
storage on disk35.5 GiBbefore W11-steady, W1126.5 GiBbefore W11-steady, W11
peak RSS35.1 GiB31.6 GiB before the writers43.6 GiB
of which anonymous4.0 GiB3.4 GiB
of which file-backed28.8 GiB34.9 GiB
disk read bytes86.0 MiB1.2 GiB
disk write bytes30.9 GiB155.5 GiB
disk read ops66,01831,170
disk write ops940,6282,647,978
syscall reads (all fds)10,524,70327,131
syscall writes (all fds)4,599,04521,040,444

measured via strawmANN: proc+cgroup / Qdrant: proc+cgroup. proc supplies syscall counts and block-layer bytes, cgroup supplies block-layer operations and bytes, so a row one interface does not carry reads n/a via …. unknown means it was not measured, and 0 means it was: an engine started with no --data-dir has no store, which is the row this section exists for. storage on disk is the level before the rows with a concurrent writer: during those an engine that rewrites segments is caught mid-rewrite, and the same row has read 3.47 and 10.11 GiB on two runs of one binary. peak RSS does include them, being a peak.

Per workload

A search row doing block-layer reads is an engine going to disk to answer a query.

workloadstrawmANN readstrawmANN writtenstrawmANN read opsstrawmANN write opsQdrant readQdrant writtenQdrant read opsQdrant write ops
W0-upload060.5 MiB0977113.7 MiB381.0 MiB3,3668,768
W000030006
W105.7 GiB011,6155.3 MiB16.1 GiB357267,470
W205.7 GiB093,568172.4 MiB25.6 GiB4,374441,697
W300030000
W400000000
W4-sat5000000000
W4-sat7000000000
W4-sat9000000000
W500000000
W6-upload188.0 KiB5.7 GiB14193,552187.3 MiB30.6 GiB4,860517,738
W600030000
W6-ef3200000000
W6-ef6400000000
W6-ef12800000000
W6-ef25600000000
W6-ef51200000000
W7-upload20.0 KiB5.7 GiB1594,166229.2 MiB30.8 GiB6,264527,400
W700000000
W8-upload100.0 KiB5.7 GiB7594,068175.3 MiB25.2 GiB4,389432,360
W800030000
W900000000
W10-ef3200000000
W10-ef6400000000
W10-ef12800000000
W10-ef25600000000
W10-ef51200000000
W12-upload588.0 KiB1.1 GiB399531,94761.1 MiB5.3 GiB1,68091,854
W12-sel10003000219
W12-sel10000000078
W12-sel1-ef3200000000
W12-sel1-ef64000000042
W12-sel1-ef128000000060
W12-sel1-ef256000000090
W12-sel1-ef512000000054
W12-sel10-ef3200000000
W12-sel10-ef6400000000
W12-sel10-ef12800000000
W12-sel10-ef25600000000
W12-sel10-ef51200000000
W1300000000
W11-steady71.8 MiB290.0 MiB55,1165,200161.2 MiB12.5 GiB3,951213,470
W1113.4 MiB1.1 GiB10,27215,52077.2 MiB8.9 GiB1,929146,672
Scheduler and memory CPU use, run-queue wait, migrations

Whether the engine was running while it ran. of wall is CPU over elapsed, waiting is runnable-but-not-scheduled, migrations checks the pinning claim, and switches gives voluntary over involuntary — the scheduler taking the core away against the engine choosing to sleep, which per unit of work is the cheapest signal of lock contention there is. A dash is not a zero: an index build's threads exit before the row does, and threads shows what fraction of the row the survivors account for.

Time spent waiting for a core

Lower is better; the bar between a pair is the gap.

Peak memory per row

Peak RSS while the row ran, from the engine's own process. Unlike the disk counters this is not refused across the two engines: residency decides where bytes live, and this is what the process held either way.

workloadstrawmANNQdrant
cpuof wallwaitingswitches vol/involmigrationsfaults min/majthreadscpuof wallwaitingswitches vol/involmigrationsfaults min/majthreads
W0-upload110.7 s542%48.2 ms47,573 / 6096113,516 / 09 1%165.4 s520%13,049.4 ms161,670 / 28,46312,323412,973 / 210,11340 7%
W07.7 s70%2.9 ms163,573 / 3700 / 0914.8 s103%3.8 ms556,507 / 14126416 / 037 0%
W111.6 s139%101.2 ms607,984 / 2701,576,446 / 09251.2 s701%106,299.2 ms158,301 / 89,63333,511666,356 / 1,598,65052 55%
W22,025.5 s769%1,210.9 ms611,814 / 9,11341,640,250 / 09 1%2,760.3 s674%117,360.9 ms181,341 / 105,71333,2361,411,741 / 1,634,90939 2%
W343.6 s92%9.6 ms189,858 / 6200 / 0980.3 s100%6.1 ms562,064 / 2881491,784 / 038 0%
W4100.7 s714%7.0 ms89,569 / 1220151 / 09114.2 s793%97,543.2 ms42,268 / 42,3804,79731,896 / 044 0%
W4-sat50190.5 s166%14.3 ms540,573 / 13403 / 09350.6 s305%12,370.5 ms1,922,913 / 9,939477,0243,772 / 051 0%
W4-sat70209.3 s255%15.0 ms485,538 / 129015 / 09379.8 s462%75,709.2 ms1,650,207 / 48,714646,3093,773 / 068
W4-sat90278.8 s436%30.3 ms440,967 / 215015 / 09417.1 s652%177,241.7 ms881,633 / 162,106340,55514,126 / 044 0%
W5100.3 s708%8.9 ms6,898 / 17800 / 0975.9 s198%12.9 ms38,987 / 4631,2892,618 / 042 0%
W6-upload2,021.2 s755%1,145.2 ms93,772 / 9,13502,205,994 / 09 0%1,196.3 s562%59,601.5 ms219,218 / 59,05032,4422,573,002 / 1,681,02841 1%
W629.6 s179%8.9 ms182,905 / 5400 / 0950.2 s203%10.9 ms539,698 / 16817,550961 / 044 0%
W6-ef3213.4 s161%2.5 ms172,962 / 600 / 0923.4 s204%7.8 ms513,951 / 6411,89998 / 044
W6-ef6420.1 s171%2.1 ms178,890 / 1200 / 0932.8 s204%7.2 ms528,971 / 9514,428357 / 044
W6-ef12829.7 s179%1.0 ms181,877 / 1800 / 0950.4 s203%11.7 ms540,542 / 16018,237584 / 044
W6-ef25650.8 s187%5.9 ms183,390 / 4504 / 0986.9 s202%36.9 ms548,858 / 34719,8521,194 / 044
W6-ef51290.5 s192%4.6 ms186,213 / 16803 / 09148.8 s201%74.8 ms556,185 / 67420,0101,936 / 044
W7-upload2,004.4 s730%1,253.2 ms99,475 / 9,39421,698,226 / 2939 0%572.4 s418%53,138.7 ms236,704 / 48,29630,4452,167,504 / 1,809,28742 0%
W720.5 s171%2.2 ms181,601 / 1800 / 0928.5 s204%9.3 ms518,417 / 5812,207841 / 046 0%
W8-upload2,850.9 s774%1,379.9 ms107,375 / 13,108141,869,849 / 09 0%6,057.1 s668%30,852.7 ms780,697 / 165,094102,0731,618,645 / 1,609,63644 1%
W863.9 s189%7.6 ms186,591 / 7500 / 0991.4 s202%79.2 ms542,438 / 36220,2444,419 / 048 0%
W9643.8 s420%102.3 ms4,276,720 / 1,23400 / 091,624.8 s798%13,447.9 ms26,815 / 23,0216,4091,478 / 050
W10-ef3232.5 s685%3.8 ms98,381 / 4400 / 0940.4 s591%15,065.5 ms247,706 / 24,60384,1211,261 / 050
W10-ef6454.5 s701%5.4 ms102,070 / 6400 / 0965.3 s633%38,874.7 ms304,258 / 42,316155,0761,194 / 067 0%
W10-ef12893.4 s707%8.8 ms107,113 / 12200 / 09107.3 s659%58,175.7 ms336,821 / 55,410178,808647 / 067
W10-ef256162.2 s706%15.2 ms113,067 / 25700 / 09186.5 s690%81,122.1 ms368,784 / 82,953193,845664 / 067
W10-ef512285.5 s704%25.5 ms116,016 / 51203 / 09335.4 s716%110,818.9 ms399,633 / 119,190195,5311,846 / 067
W12-upload297.7 s650%166.6 ms25,385 / 1,4680445,699 / 349 1%520.5 s528%10,982.4 ms67,814 / 12,4057,863372,899 / 446,81551 0%
W12-sel149.5 s185%4.5 ms126,934 / 4100 / 0962.7 s202%21.6 ms533,653 / 14917,4791,065 / 056
W12-sel10118.1 s194%6.9 ms184,898 / 30300 / 09122.3 s201%131.2 ms538,093 / 37419,113997 / 056
W12-sel1-ef3249.5 s184%1.1 ms125,326 / 2600 / 0929.8 s202%11.6 ms511,180 / 3810,549304 / 056
W12-sel1-ef6449.4 s184%1.7 ms120,799 / 3100 / 0942.0 s202%10.9 ms526,563 / 7913,463309 / 056
W12-sel1-ef12849.5 s184%1.4 ms121,877 / 3000 / 0961.8 s202%19.6 ms535,564 / 22117,422581 / 056
W12-sel1-ef25649.5 s184%1.1 ms126,169 / 3000 / 0993.2 s202%118.2 ms525,768 / 36118,434723 / 056
W12-sel1-ef51249.4 s184%1.9 ms129,987 / 3700 / 09131.4 s199%226.9 ms513,873 / 65516,5251,210 / 056
W12-sel10-ef3247.0 s185%7.4 ms178,512 / 3600 / 0944.9 s202%10.7 ms534,379 / 14515,993652 / 056
W12-sel10-ef6472.6 s190%10.3 ms181,617 / 8300 / 0972.3 s202%38.0 ms538,686 / 31819,122736 / 056
W12-sel10-ef128118.0 s194%10.3 ms185,057 / 28200 / 09122.5 s201%154.7 ms538,246 / 55219,6631,051 / 056
W12-sel10-ef256201.5 s193%17.9 ms172,673 / 63700 / 09205.8 s200%163.6 ms562,966 / 1,61422,1762,134 / 056
W12-sel10-ef512325.0 s191%38.1 ms125,238 / 1,16900 / 09348.9 s200%140.4 ms570,358 / 2,42823,5634,057 / 056
W135.6 s138%4.2 ms242,724 / 1300 / 099.1 s413%853.7 ms381,104 / 4,420108,063780 / 066
W11-steady351.5 s781%263,490.0 ms80,207 / 116,1011,118580,619 / 91,2219 48%900.4 s787%784,454.5 ms244,381 / 282,24458,423889,903 / 147,52059 0%
W11476.4 s1,503%208,739.9 ms28,280 / 81,7338061,020,754 / 298,41018 27%652.0 s813%518,000.9 ms124,645 / 177,33825,478769,832 / 324,97066 85%
Waiting and stalls where the time went when the engine was not running

waiting is runnable and not scheduled, blocked on disk is not runnable at all, and pressure is PSI for the engine's cgroup (some = at least one task stalled, full = every runnable task).

Time stalled on I/O

Time every task in the engine's cgroup was stalled on I/O. Lower is better. 34 rows never stalled on I/O and are not drawn.

Time stalled on memory

Time every task in the engine's cgroup was stalled on memory. Lower is better. 40 rows never stalled on memory and are not drawn.

workloadstrawmANNQdrant
waitingblocked on diskcpu some/fullio some/fullmemory some/fullwaitingblocked on diskcpu some/fullio some/fullmemory some/full
W0-upload48.2 ms0 ms14 / 10 ms1 / 1 ms0 / 0 ms13,049.4 ms0 ms1,202 / 18 ms69 / 66 ms0 / 0 ms
W02.9 ms0 ms47 / 47 ms0 / 0 ms0 / 0 ms3.8 ms0 ms78 / 78 ms0 / 0 ms0 / 0 ms
W1101.2 ms0 ms55 / 55 ms2 / 2 ms0 / 0 ms106,299.2 ms0 ms9,340 / 29 ms859 / 329 ms0 / 0 ms
W21,210.9 ms0 ms190 / 74 ms1 / 1 ms0 / 0 ms117,360.9 ms0 ms10,277 / 79 ms4,954 / 4,332 ms0 / 0 ms
W39.6 ms0 ms40 / 40 ms0 / 0 ms0 / 0 ms6.1 ms0 ms108 / 107 ms0 / 0 ms0 / 0 ms
W47.0 ms0 ms3 / 3 ms0 / 0 ms0 / 0 ms97,543.2 ms0 ms5,034 / 2 ms0 / 0 ms0 / 0 ms
W4-sat5014.3 ms0 ms156 / 156 ms0 / 0 ms0 / 0 ms12,370.5 ms0 ms1,728 / 270 ms0 / 0 ms0 / 0 ms
W4-sat7015.0 ms0 ms156 / 156 ms0 / 0 ms0 / 0 ms75,709.2 ms0 ms8,347 / 206 ms0 / 0 ms0 / 0 ms
W4-sat9030.3 ms0 ms139 / 139 ms0 / 0 ms0 / 0 ms177,241.7 ms0 ms16,413 / 93 ms0 / 0 ms0 / 0 ms
W58.9 ms0 ms2 / 2 ms0 / 0 ms0 / 0 ms12.9 ms0 ms15 / 15 ms0 / 0 ms0 / 0 ms
W6-upload1,145.2 ms0 ms156 / 44 ms13 / 13 ms0 / 0 ms59,601.5 ms0 ms5,586 / 77 ms10,395 / 9,918 ms0 / 0 ms
W68.9 ms0 ms36 / 36 ms0 / 0 ms0 / 0 ms10.9 ms0 ms67 / 66 ms0 / 0 ms0 / 0 ms
W6-ef322.5 ms0 ms31 / 31 ms0 / 0 ms0 / 0 ms7.8 ms0 ms55 / 54 ms0 / 0 ms0 / 0 ms
W6-ef642.1 ms0 ms38 / 38 ms0 / 0 ms0 / 0 ms7.2 ms0 ms60 / 59 ms0 / 0 ms0 / 0 ms
W6-ef1281.0 ms0 ms36 / 36 ms0 / 0 ms0 / 0 ms11.7 ms0 ms68 / 67 ms0 / 0 ms0 / 0 ms
W6-ef2565.9 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms36.9 ms0 ms116 / 113 ms0 / 0 ms0 / 0 ms
W6-ef5124.6 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms74.8 ms0 ms148 / 142 ms0 / 0 ms0 / 0 ms
W7-upload1,253.2 ms0 ms184 / 55 ms42 / 42 ms563 / 563 ms53,138.7 ms0 ms4,999 / 88 ms9,288 / 7,514 ms0 / 0 ms
W72.2 ms0 ms34 / 34 ms0 / 0 ms0 / 0 ms9.3 ms0 ms57 / 57 ms0 / 0 ms0 / 0 ms
W8-upload1,379.9 ms0 ms189 / 48 ms62 / 62 ms84 / 84 ms30,852.7 ms0 ms7,227 / 298 ms3,991 / 2,329 ms0 / 0 ms
W87.6 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms79.2 ms0 ms120 / 111 ms0 / 0 ms0 / 0 ms
W9102.3 ms0 ms505 / 505 ms0 / 0 ms0 / 0 ms13,447.9 ms0 ms1,617 / 38 ms0 / 0 ms0 / 0 ms
W10-ef323.8 ms0 ms10 / 10 ms0 / 0 ms0 / 0 ms15,065.5 ms0 ms1,768 / 23 ms0 / 0 ms0 / 0 ms
W10-ef645.4 ms0 ms9 / 9 ms0 / 0 ms0 / 0 ms38,874.7 ms0 ms3,754 / 23 ms0 / 0 ms0 / 0 ms
W10-ef1288.8 ms0 ms9 / 9 ms0 / 0 ms0 / 0 ms58,175.7 ms0 ms5,650 / 34 ms0 / 0 ms0 / 0 ms
W10-ef25615.2 ms0 ms8 / 8 ms0 / 0 ms0 / 0 ms81,122.1 ms0 ms8,322 / 35 ms0 / 0 ms0 / 0 ms
W10-ef51225.5 ms0 ms7 / 7 ms0 / 0 ms0 / 0 ms110,818.9 ms0 ms11,792 / 64 ms0 / 0 ms0 / 0 ms
W12-upload166.6 ms0 ms25 / 9 ms54 / 54 ms2,219 / 2,219 ms10,982.4 ms0 ms1,054 / 29 ms266 / 171 ms0 / 0 ms
W12-sel14.5 ms0 ms35 / 35 ms0 / 0 ms0 / 0 ms21.6 ms0 ms78 / 76 ms0 / 0 ms0 / 0 ms
W12-sel106.9 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms131.2 ms0 ms143 / 128 ms0 / 0 ms0 / 0 ms
W12-sel1-ef321.1 ms0 ms35 / 35 ms0 / 0 ms0 / 0 ms11.6 ms0 ms58 / 57 ms0 / 0 ms0 / 0 ms
W12-sel1-ef641.7 ms0 ms35 / 35 ms0 / 0 ms0 / 0 ms10.9 ms0 ms63 / 62 ms0 / 0 ms0 / 0 ms
W12-sel1-ef1281.4 ms0 ms35 / 35 ms0 / 0 ms0 / 0 ms19.6 ms0 ms76 / 75 ms0 / 0 ms0 / 0 ms
W12-sel1-ef2561.1 ms0 ms35 / 35 ms0 / 0 ms0 / 0 ms118.2 ms0 ms125 / 111 ms0 / 0 ms0 / 0 ms
W12-sel1-ef5121.9 ms0 ms35 / 35 ms0 / 0 ms0 / 0 ms226.9 ms0 ms145 / 122 ms0 / 0 ms0 / 0 ms
W12-sel10-ef327.4 ms0 ms36 / 36 ms0 / 0 ms0 / 0 ms10.7 ms0 ms65 / 64 ms0 / 0 ms0 / 0 ms
W12-sel10-ef6410.3 ms0 ms36 / 36 ms0 / 0 ms0 / 0 ms38.0 ms0 ms94 / 90 ms0 / 0 ms0 / 0 ms
W12-sel10-ef12810.3 ms0 ms38 / 38 ms0 / 0 ms0 / 0 ms154.7 ms0 ms147 / 129 ms0 / 0 ms0 / 0 ms
W12-sel10-ef25617.9 ms0 ms41 / 41 ms0 / 0 ms0 / 0 ms163.6 ms0 ms166 / 152 ms0 / 0 ms0 / 0 ms
W12-sel10-ef51238.1 ms0 ms42 / 42 ms0 / 0 ms0 / 0 ms140.4 ms0 ms202 / 191 ms0 / 0 ms0 / 0 ms
W134.2 ms0 ms28 / 28 ms0 / 0 ms0 / 0 ms853.7 ms0 ms121 / 26 ms0 / 0 ms0 / 0 ms
W11-steady263,490.0 ms0 ms33,892 / 11 ms490 / 90 ms0 / 0 ms784,454.5 ms0 ms85,464 / 64 ms1,684 / 46 ms0 / 0 ms
W11208,739.9 ms0 ms26,851 / 13 ms100 / 71 ms0 / 0 ms518,000.9 ms0 ms52,090 / 34 ms1,886 / 131 ms0 / 0 ms

A dash under blocked on disk is not a zero: most hosts ship with kernel.task_delayacct=0, which reports the field as a permanent zero and would manufacture the strongest claim here — that the memory-resident engine never waits for a device — out of a sysctl. bench/setup.py apply turns it on.

Hardware counters cycles, DRAM loads, TLB walks, IPC per query

What the core did per query, from perf stat attached to the engine for the same window as the /proc counters. §5's cost model is a set of claims about instructions per cycle, DRAM traffic and TLB reach; these are those quantities, measured on the row the headline quotes rather than on a microbenchmark.

What one query cost, in cycles

Lower is better; the bar between a pair is the gap. Load rows are absent: they have no queries to divide by, and an absolute count under a per-query axis would be a different quantity wearing this one's label.

Demand loads from DRAM per query

Demand loads served from DRAM, times the cache line this host reports. Lower is better. Demand only: the hardware prefetcher's fills are a separate counter and are not in this number, so a row that streams — an exact scan above all — moved far more than this says. On the graph rows, where there is little for a prefetcher to predict, it is most of the traffic and is the outstanding-miss quantity §5.2 argues the engine is limited by.

TLB walks per query

Data TLB misses that reached a page walk, per query. Lower is better. §5.5 argues a memory-resident index needs hugepage care; this is what not taking it costs, and `--no-huge-pages` is the A/B arm that prices it.

Instructions per cycle

How well the core was fed while it ran. Not a score: a low IPC on a memory-bound row is what §5.2 predicts, and a high one on a row that does more work per query is not a win. Read it beside the two charts above.

Branch mispredictions per 1k instructions

A graph traversal is a chain of data-dependent branches and none of them is predictable from the last query, so this is the column that separates a slow kernel from an unpredictable walk. Per thousand instructions rather than per branch: `branches` costs a PMU counter that `ref-cycles` needs more, and MPKI is the comparable form anyway.

workloadstrawmANNQdrant
IPCGHzbranch MPKIcycles/querydemand DRAM/queryTLB walks/queryIPCGHzbranch MPKIcycles/querydemand DRAM/queryTLB walks/query
W0-upload1.502.001.004x nominal5.36---1.302.001.004x nominal4.53---
W01.262.011.005x nominal5.83253,1186.8 KiB987.61.172.011.005x nominal3.07462,93137.6 KiB1,253.1
W12.012.011.005x nominal0.69---1.102.011.005x nominal1.23---
W20.422.011.005x nominal2.35---0.602.011.005x nominal1.74---
W30.712.011.006x nominal3.841,673,9085,795.5 KiB4,265.80.682.011.005x nominal1.933,070,5974,539.1 KiB6,973.2
W40.292.011.005x nominal3.244,018,2974,318.8 KiB4,073.20.442.001.004x nominal2.004,515,4494,229.3 KiB6,179.4
W4-sat500.642.011.006x nominal3.861,852,6365,434.5 KiB4,209.30.612.011.005x nominal2.183,399,2024,503.4 KiB6,939.2
W4-sat700.582.011.005x nominal3.892,040,1795,093.9 KiB4,182.90.572.011.005x nominal2.263,646,9474,415.5 KiB6,865.3
W4-sat900.442.011.005x nominal3.762,718,3874,460.7 KiB4,161.30.502.011.006x nominal2.234,096,7394,312.2 KiB6,567.0
W50.282.011.005x nominal3.454,012,8544,352.8 KiB4,013.30.602.011.005x nominal2.183,021,5954,680.5 KiB5,890.6
W6-upload0.422.011.005x nominal2.28---1.682.011.005x nominal1.48---
W60.892.011.006x nominal4.961,119,1173,024.8 KiB6,182.61.502.011.006x nominal1.611,881,039747.9 KiB4,767.7
W6-ef320.872.001.004x nominal3.57480,384956.7 KiB2,390.11.602.011.007x nominal1.27818,939242.0 KiB2,165.5
W6-ef640.842.011.006x nominal4.41735,2591,678.6 KiB3,768.91.552.011.006x nominal1.441,189,321417.7 KiB3,102.5
W6-ef1280.892.011.006x nominal4.991,120,5823,027.8 KiB6,191.41.492.011.006x nominal1.651,888,084746.6 KiB4,802.1
W6-ef2560.882.011.006x nominal5.481,965,6845,423.8 KiB11,026.41.392.011.005x nominal2.003,299,7601,439.3 KiB7,732.8
W6-ef5120.882.011.006x nominal5.933,558,5289,787.3 KiB21,026.71.362.011.006x nominal2.135,749,6202,682.8 KiB12,895.3
W7-upload0.422.011.005x nominal2.76---2.042.001.004x nominal4.72---
W71.082.011.005x nominal7.99762,362447.3 KiB3,505.11.662.011.005x nominal3.211,018,381350.6 KiB3,853.2
W8-upload2.412.011.005x nominal0.73---5.402.001.004x nominal0.17---
W84.072.011.006x nominal0.592,489,262469.4 KiB4,458.32.952.011.006x nominal0.623,483,686508.3 KiB5,167.2
W90.622.001.004x nominal0.33621,894,169752,253.1 KiB1,472,011.80.362.011.005x nominal0.091,627,503,846525,904.2 KiB1,421,558.3
W10-ef320.372.011.005x nominal2.411,281,2371,354.1 KiB2,268.70.632.011.005x nominal1.951,538,4531,527.7 KiB2,637.3
W10-ef640.342.011.005x nominal2.932,161,4722,314.8 KiB3,741.60.562.011.005x nominal2.142,478,1032,541.6 KiB4,045.6
W10-ef1280.322.011.005x nominal3.313,716,0244,011.3 KiB6,516.40.512.001.004x nominal2.334,114,7184,308.2 KiB6,472.2
W10-ef2560.312.011.005x nominal3.636,475,3556,999.6 KiB11,791.50.472.001.004x nominal2.467,237,9237,521.7 KiB10,654.9
W10-ef5120.322.011.005x nominal4.0011,425,81312,417.5 KiB21,944.00.432.001.004x nominal2.8013,132,73813,522.6 KiB18,055.8
W12-upload0.502.011.005x nominal2.46---0.852.011.005x nominal1.98---
W12-sel10.502.011.005x nominal2.921,927,7387,189.7 KiB4,611.60.842.001.004x nominal3.142,364,7692,530.8 KiB4,490.1
W12-sel101.022.011.005x nominal3.964,660,2129,143.5 KiB10,406.00.822.011.005x nominal2.314,702,4005,188.2 KiB7,838.2
W12-sel1-ef320.502.011.005x nominal2.911,926,6457,193.4 KiB4,615.20.942.011.005x nominal2.181,076,3781,080.2 KiB2,432.0
W12-sel1-ef640.502.011.005x nominal2.931,927,4277,195.5 KiB4,617.00.892.011.005x nominal2.551,558,7311,667.3 KiB3,318.2
W12-sel1-ef1280.502.011.005x nominal2.951,926,3387,195.6 KiB4,615.30.852.011.006x nominal3.102,339,8762,512.8 KiB4,491.9
W12-sel1-ef2560.502.011.006x nominal2.931,929,6687,193.5 KiB4,611.30.832.011.005x nominal3.833,552,0563,599.0 KiB5,763.9
W12-sel1-ef5120.502.011.006x nominal2.921,929,4377,189.5 KiB4,616.90.892.011.006x nominal4.165,083,3854,653.3 KiB7,000.2
W12-sel10-ef320.992.011.006x nominal3.241,819,2193,365.4 KiB3,875.10.952.011.005x nominal1.781,667,6591,688.1 KiB3,336.8
W12-sel10-ef640.992.011.005x nominal3.672,841,1065,584.0 KiB6,118.50.862.011.005x nominal2.082,733,7042,933.6 KiB4,987.1
W12-sel10-ef1281.022.011.006x nominal3.974,664,6929,152.2 KiB10,396.90.822.011.005x nominal2.314,703,6265,187.6 KiB7,847.4
W12-sel10-ef2561.052.011.006x nominal4.188,001,69814,867.8 KiB18,748.60.822.011.005x nominal2.388,011,9308,876.0 KiB12,544.3
W12-sel10-ef5120.602.011.005x nominal0.7712,956,69341,598.7 KiB33,839.10.832.001.004x nominal2.5313,720,88714,608.6 KiB19,692.4
W131.362.011.007x nominal1.2042,2841.0 KiB12.71.472.011.006x nominal1.41137,3521.4 KiB40.0
W11-steady0.332.001.005x nominal1.8722,322,25326,845.9 KiB29,355.20.522.011.005x nominal0.5657,372,41827,374.6 KiB56,061.9
W110.372.001.005x nominal1.64119,356,527157,005.0 KiB155,746.30.572.011.005x nominal0.53162,278,41883,060.1 KiB159,785.1

Nothing here is scaled. Where perf had to multiplex the group, the values it prints are extrapolations from the fraction of the row each counter was on, and they are withheld rather than shown — the same rule the scheduler table applies to a thread that exited. An event this host does not implement is likewise blank, never zero.

The two instruments disagree on these rows. strawmANN: perf and /proc disagree on minor faults by up to 13% on W12-upload; strawmANN: perf and /proc disagree on context switches by up to 9% on W11-steady, W11; Qdrant: perf and /proc disagree on minor faults by up to 43% on W1, W2, W6-upload, W7-upload, W8-upload, W12-upload, W11-steady, W11; Qdrant: perf and /proc disagree on context switches by up to 65% on W7-upload, W8-upload, W12-upload. perf keeps an exiting thread's counts and the /proc sums do not, so a gap here is usually the same lost-thread effect the scheduler table reports as coverage.

Collection config what each engine says it built

Read back from the engine rather than taken from the request. --segments 1 sets Qdrant's default_segment_number, which its optimizer treats as a target, and the segment count is the largest confound in the ef comparison. strawmANN held 1 segment; Qdrant held 2 segments.

Where the engines disagree a cell reads strawmANN / Qdrant, and is marked.

collectionsegmentspopulated segmentsrequested segmentspointsindexed vectorsvector sizehnsw mhnsw ef_constructquantizationshardsvector residency
bench01 / 2- / 1- / 1990,000990,000416100-1cached
bench11 / 5- / 5- / 1990,0000 / 57,2001,53616100-1cached
bench21 / 2- / 1- / 1990,000990,0001,53616100-1cached
bench61 / 2- / 1- / 1990,000990,0001,53616100scalar1cached
bench71 / 2- / 1- / 1990,000990,0001,53616100binary1cached
bench81 / 2- / 1- / 1990,000990,0001,53616100product1cached
bench121 / 2- / 1- / 1200,000200,0001,53616100-1cached

After the mutating rows bench2 held strawmANN 1,237,500 (+247,500), 1,137,262 of them indexed; Qdrant 1,237,500 (+247,500), 1,040,500 of them indexed. The table above is the state the throughput, latency and recall rows searched; this is what W11 left, and is that row's subject rather than theirs.

bench1 was read back straight after its row and dropped, with no settle between, so an engine still building it reports a snapshot taken mid-ingest. Its cells are not marked as a disagreement.

Per workload every metric of every row, engine beside engine

Metric down, engine across, so a comparison is two adjacent cells. Metrics a row did not measure are dropped rather than shown empty.

W0-upload transport floor: load

upload · collection=bench0 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=990,000
metricstrawmANNQdrant
wall clock20.4 s31.8 s1.56x
of which upload1.3 s7.5 s5.80x
of which index wait17.1 s22.1 s1.29x
time to Green17.1 s22.1 s1.29x
cpu110.7 s165.4 s1.49x
cpu, % of wall542%520%
waiting for a core48.2 ms13,049.4 ms270.66x
migrations612,3232,053.83x
faults min/maj113,516 / 0412,973 / 210,113
peak RSS440.5 MiB1.6 GiB3.65x
disk written60.5 MiB381.0 MiB
disk read0113.7 MiB

W0 d=4 floor: graph traversal with the distance taken out 1.32x

closed-loop · collection=bench0 · client -p 1 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second4,5983,4951.32x
client p50213 µs285 µs1.34x
client p99261 µs348 µs1.34x
client p99.9305 µs400 µs1.31x
server p50121 µs182 µs1.50x
server p99133 µs218 µs1.65x
wall clock10.9 s14.3 s1.31x
queries sent50,00050,000
cpu7.7 s14.8 s1.92x
cpu, % of wall70%103%
waiting for a core2.9 ms3.8 ms1.30x
migrations0126
faults min/maj0 / 0416 / 0
peak RSS440.5 MiB1.6 GiB3.65x
disk written00
disk read00
1.32x is 1.83x less work per query x 0.67x cores busy during the row (0.67 against 1.01) x 1.00x clock x 1.08x counted on-CPU share: the middle term is occupancy, not search speed d=4 floor: this measures graph traversal and heaps with the distance taken out, not request plumbing (findings 28: 76% of the server's CPU was HNSW search, 24% kernel; the RPC path is a small remainder)

W1 ingest throughput

upload · collection=bench1 · client -p 8 -t 8 -c 1 (-c bfb default) · n=990,000
metricstrawmANNQdrant
wall clock8.4 s35.8 s4.26x
of which upload5.5 s33.6 s6.07x
cpu11.6 s251.2 s21.58x
cpu, % of wall139%701%
waiting for a core101.2 ms106,299.2 ms1,050.45x
migrations033,511
faults min/maj1,576,446 / 0666,356 / 1,598,650
peak RSS7.6 GiB12.3 GiB1.61x
disk written5.7 GiB16.1 GiB
disk read05.3 MiB

W2 index build time

upload · collection=bench2 · client -p 8 -t 8 -c 1 (-c bfb default) · n=990,000
metricstrawmANNQdrant
wall clock263.2 s409.6 s1.56x
of which upload5.5 s35.8 s6.49x
of which index wait254.9 s370.9 s1.46x
time to Green254.9 s370.9 s1.46x
cpu2,025.5 s2,760.3 s1.36x
cpu, % of wall769%674%
waiting for a core1,210.9 ms117,360.9 ms96.92x
migrations433,2368,309.00x
faults min/maj1,640,250 / 01,411,741 / 1,634,909
peak RSS7.9 GiB18.1 GiB2.30x
disk written5.7 GiB25.6 GiB
disk read0172.4 MiB

W3 search, fp32, single query 1.70x

ef=128 · closed-loop · collection=bench2 · client -p 1 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,0606241.70x
recall@100.96670.9694~equal
client p50947 µs1.60 ms1.69x
client p991.37 ms2.29 ms1.67x
client p99.91.89 ms2.57 ms1.36x
server p50845 µs1.47 ms1.74x
server p991.26 ms2.15 ms1.71x
wall clock47.2 s80.2 s1.70x
queries sent50,00050,000
cpu43.6 s80.3 s1.84x
cpu, % of wall92%100%
waiting for a core9.6 ms6.1 ms0.63x
migrations0149
faults min/maj0 / 01,784 / 0
peak RSS7.9 GiB18.1 GiB2.30x
disk written00
disk read00

W4 search, saturating (closed loop) 1.02x

ef=128 · closed-loop · collection=bench2 · client -p 64 -t 16 -c 2 · n=50,000
metricstrawmANNQdrant
queries/second3,5563,483~equal
recall@100.96670.9694~equal
client p5018.05 ms18.14 ms~equal
client p9920.22 ms32.49 ms1.61x
client p99.926.83 ms43.64 ms1.63x
server p5017.93 ms17.64 ms~equal
server p9920.07 ms27.56 ms1.37x
wall clock14.1 s14.4 s1.02x
queries sent50,00050,000
cpu100.7 s114.2 s1.13x
cpu, % of wall714%793%
waiting for a core7.0 ms97,543.2 ms13,948.10x
migrations04,797
faults min/maj151 / 031,896 / 0
peak RSS7.9 GiB18.1 GiB2.30x
disk written00
disk read00

W4-sat50 search, fixed rate at 50% of saturation (open loop) offered

ef=128 · open-loop · offered=1,741/s (50% of saturation, pinned reference 3,481 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second1,7411,741offered
recall@100.96670.9694~equal
client p501.03 ms1.76 ms1.72x
client p991.63 ms2.95 ms1.81x
client p99.91.87 ms3.77 ms2.01x
server p50922 µs1.61 ms1.75x
server p991.52 ms2.73 ms1.79x
wall clock115.0 s115.0 s~equal
queries sent200,000200,000
cpu190.5 s350.6 s1.84x
cpu, % of wall166%305%
waiting for a core14.3 ms12,370.5 ms868.04x
migrations0477,024
faults min/maj3 / 03,772 / 0
peak RSS7.9 GiB18.1 GiB2.30x
disk written00
disk read00

W4-sat70 search, fixed rate at 70% of saturation (open loop) offered

ef=128 · open-loop · offered=2,437/s (70% of saturation, pinned reference 3,481 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second2,4372,437offered
recall@100.96670.9694~equal
client p501.11 ms2.00 ms1.80x
client p991.95 ms3.75 ms1.93x
client p99.92.29 ms4.76 ms2.08x
server p501.01 ms1.82 ms1.81x
server p991.84 ms3.45 ms1.88x
wall clock82.2 s82.2 s~equal
queries sent200,000200,000
cpu209.3 s379.8 s1.81x
cpu, % of wall255%462%
waiting for a core15.0 ms75,709.2 ms5,060.41x
migrations0646,309
faults min/maj15 / 03,773 / 0
peak RSS7.9 GiB18.1 GiB2.30x
disk written00
disk read00

W4-sat90 search, fixed rate at 90% of saturation (open loop) offered

ef=128 · open-loop · offered=3,133/s (90% of saturation, pinned reference 3,481 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second3,1333,133offered
recall@100.96670.9694~equal
client p501.39 ms2.70 ms1.94x
client p992.83 ms9.53 ms3.37x
client p99.93.56 ms17.91 ms5.03x
server p501.28 ms2.38 ms1.86x
server p992.70 ms7.12 ms2.64x
wall clock64.0 s64.0 s~equal
queries sent200,000200,000
cpu278.8 s417.1 s1.50x
cpu, % of wall436%652%
waiting for a core30.3 ms177,241.7 ms5,847.05x
migrations0340,555
faults min/maj15 / 014,126 / 0
peak RSS7.9 GiB18.1 GiB2.30x
disk written00
disk read00

W5 search batched (16 distinct dataset queries per request) 2.71x

ef=128 · closed-loop · collection=bench2 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second3,5301,3032.71x
recall@100.96670.9694~equal
client p509.06 ms24.54 ms2.71x
client p9910.55 ms27.79 ms2.63x
client p99.911.31 ms28.86 ms2.55x
server p508.41 ms23.68 ms2.81x
server p999.90 ms26.88 ms2.72x
wall clock14.2 s38.4 s2.71x
queries sent50,00050,000
cpu100.3 s75.9 s0.76x
cpu, % of wall708%198%
waiting for a core8.9 ms12.9 ms1.44x
migrations01,289
faults min/maj0 / 02,618 / 0
peak RSS7.9 GiB18.1 GiB2.30x
disk written00
disk read00
per-batch latency (16 queries/request); queries: dataset, random-sample 2.71x is 0.75x less work per query x 3.58x cores busy during the row (7.07 against 1.97) x 1.00x clock: the middle term is occupancy, not search speed batched 16, so qps and rps differ by 16x; the ratio uses qps

W6-upload scalar quantization: load

upload · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=990,000
metricstrawmANNQdrant
wall clock267.6 s212.8 s0.80x
of which upload5.8 s31.2 s5.37x
of which index wait258.9 s176.4 s0.68x
time to Green258.9 s176.4 s0.68x
cpu2,021.2 s1,196.3 s0.59x
cpu, % of wall755%562%
waiting for a core1,145.2 ms59,601.5 ms52.04x
migrations032,442
faults min/maj2,205,994 / 02,573,002 / 1,681,028
peak RSS17.4 GiB32.3 GiB1.86x
disk written5.7 GiB30.6 GiB
disk read188.0 KiB187.3 MiB

W6 quantized: scalar 1.50x

ef=128 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second3,0282,0261.50x
recall@100.96690.9682~equal
client p50663 µs987 µs1.49x
client p99867 µs1.30 ms1.50x
client p99.91.00 ms1.43 ms1.43x
server p50568 µs872 µs1.54x
server p99766 µs1.18 ms1.53x
wall clock16.6 s24.7 s1.49x
queries sent50,00050,000
cpu29.6 s50.2 s1.70x
cpu, % of wall179%203%
waiting for a core8.9 ms10.9 ms1.22x
migrations017,550
faults min/maj0 / 0961 / 0
peak RSS17.4 GiB32.3 GiB1.86x
disk written00
disk read00

W6-ef32 SQ8 recall control, ef=32 (latency only) 1.39x

ef=32 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=3.2 · rescore=true
metricstrawmANNQdrant
queries/second6,0514,3611.39x
recall@100.89180.8906~equal
client p50313 µs454 µs1.45x
client p99587 µs617 µs1.05x
client p99.9705 µs715 µs~equal
server p50223 µs347 µs1.56x
server p99475 µs508 µs1.07x
wall clock8.3 s11.5 s1.39x
queries sent50,00050,000
cpu13.4 s23.4 s1.75x
cpu, % of wall161%204%
waiting for a core2.5 ms7.8 ms3.11x
migrations011,899
faults min/maj0 / 098 / 0
peak RSS17.4 GiB32.3 GiB1.86x
disk written00
disk read00
1.39x is 1.70x less work per query x 0.79x cores busy during the row (1.59 against 2.02) x 1.00x clock x 1.04x counted on-CPU share: the middle term is occupancy, not search speed

W6-ef64 SQ8 recall control, ef=64 (latency only) 1.37x

ef=64 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=6.4 · rescore=true
metricstrawmANNQdrant
queries/second4,2663,1091.37x
recall@100.93930.9413~equal
client p50458 µs642 µs1.40x
client p99826 µs842 µs~equal
client p99.9981 µs998 µs~equal
server p50364 µs531 µs1.46x
server p99712 µs727 µs1.02x
wall clock11.8 s16.1 s1.37x
queries sent50,00050,000
cpu20.1 s32.8 s1.64x
cpu, % of wall171%204%
waiting for a core2.1 ms7.2 ms3.38x
migrations014,428
faults min/maj0 / 0357 / 0
peak RSS17.4 GiB32.3 GiB1.86x
disk written00
disk read00

W6-ef128 SQ8 recall control, ef=128 (latency only) 1.50x

ef=128 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second3,0222,0181.50x
recall@100.96690.9682~equal
client p50665 µs991 µs1.49x
client p99866 µs1.30 ms1.50x
client p99.9984 µs1.43 ms1.46x
server p50569 µs875 µs1.54x
server p99766 µs1.18 ms1.54x
wall clock16.6 s24.8 s1.50x
queries sent50,00050,000
cpu29.7 s50.4 s1.70x
cpu, % of wall179%203%
waiting for a core1.0 ms11.7 ms11.36x
migrations018,237
faults min/maj0 / 0584 / 0
peak RSS17.4 GiB32.3 GiB1.86x
disk written00
disk read00

W6-ef256 SQ8 recall control, ef=256 (latency only) 1.58x

ef=256 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=25.6 · rescore=true
metricstrawmANNQdrant
queries/second1,8411,1641.58x
recall@100.98110.9820~equal
client p501.09 ms1.71 ms1.57x
client p991.41 ms2.31 ms1.64x
client p99.91.56 ms2.45 ms1.57x
server p50989 µs1.58 ms1.60x
server p991.31 ms2.18 ms1.66x
wall clock27.2 s43.0 s1.58x
queries sent50,00050,000
cpu50.8 s86.9 s1.71x
cpu, % of wall187%202%
waiting for a core5.9 ms36.9 ms6.24x
migrations019,852
faults min/maj4 / 01,194 / 0
peak RSS17.4 GiB32.3 GiB1.86x
disk written00
disk read00

W6-ef512 SQ8 recall control, ef=512 (latency only) 1.58x

ef=512 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=51.2 · rescore=true
metricstrawmANNQdrant
queries/second1,0636741.58x
recall@100.98950.9900~equal
client p501.88 ms2.95 ms1.57x
client p992.48 ms4.06 ms1.63x
client p99.92.63 ms4.25 ms1.62x
server p501.78 ms2.80 ms1.57x
server p992.38 ms3.90 ms1.64x
wall clock47.1 s74.2 s1.58x
queries sent50,00050,000
cpu90.5 s148.8 s1.64x
cpu, % of wall192%201%
waiting for a core4.6 ms74.8 ms16.43x
migrations020,010
faults min/maj3 / 01,936 / 0
peak RSS17.4 GiB32.3 GiB1.86x
disk written00
disk read00

W7-upload binary quantization: load

upload · collection=bench7 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=990,000
metricstrawmANNQdrant
wall clock274.7 s136.9 s0.50x
of which upload5.9 s45.4 s7.66x
of which index wait263.9 s89.2 s0.34x
time to Green263.9 s89.2 s0.34x
cpu2,004.4 s572.4 s0.29x
cpu, % of wall730%418%
waiting for a core1,253.2 ms53,138.7 ms42.40x
migrations230,44515,222.50x
faults min/maj1,698,226 / 2932,167,504 / 1,809,287
peak RSS23.6 GiB37.1 GiB1.57x
disk written5.7 GiB30.8 GiB
disk read20.0 KiB229.2 MiB

W7 quantized: binary + oversampling 1.16x

ef=128 · closed-loop · collection=bench7 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second4,1683,5791.16x
recall@100.94700.9456~equal
client p50459 µs558 µs1.21x
client p99670 µs671 µs~equal
client p99.9767 µs777 µs~equal
server p50364 µs450 µs1.23x
server p99556 µs553 µs~equal
wall clock12.0 s14.0 s1.16x
queries sent50,00050,000
cpu20.5 s28.5 s1.39x
cpu, % of wall171%204%
waiting for a core2.2 ms9.3 ms4.22x
migrations012,207
faults min/maj0 / 0841 / 0
peak RSS23.6 GiB37.1 GiB1.57x
disk written00
disk read00

W8-upload PQ: load

upload · collection=bench8 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=990,000
metricstrawmANNQdrant
wall clock368.5 s907.0 s2.46x
of which upload6.7 s28.7 s4.30x
of which index wait359.3 s875.1 s2.44x
time to Green359.3 s875.1 s2.44x
cpu2,850.9 s6,057.1 s2.12x
cpu, % of wall774%668%
waiting for a core1,379.9 ms30,852.7 ms22.36x
migrations14102,0737,290.93x
faults min/maj1,869,849 / 01,618,645 / 1,609,636
peak RSS25.0 GiB43.6 GiB1.74x
disk written5.7 GiB25.2 GiB
disk read100.0 KiB175.3 MiB

W8 quantized: PQ 1.34x

ef=128 · closed-loop · collection=bench8 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second1,4821,1061.34x
recall@100.96540.9644~equal
client p501.36 ms1.81 ms1.33x
client p991.77 ms2.32 ms1.31x
client p99.91.99 ms2.60 ms1.30x
server p501.26 ms1.68 ms1.34x
server p991.66 ms2.17 ms1.31x
wall clock33.8 s45.3 s1.34x
queries sent50,00050,000
cpu63.9 s91.4 s1.43x
cpu, % of wall189%202%
waiting for a core7.6 ms79.2 ms10.46x
migrations020,244
faults min/maj0 / 04,419 / 0
peak RSS25.0 GiB43.6 GiB1.74x
disk written00
disk read00

W9 exact / brute force 1.33x

exact · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=2,000
metricstrawmANNQdrant
queries/second13101.33x
client p50575.52 ms808.83 ms1.41x
client p991,504.90 ms871.60 ms0.58x
client p99.92,481.43 ms886.78 ms0.36x
server p50574.47 ms807.41 ms1.41x
server p991,504.22 ms869.93 ms0.58x
wall clock153.1 s203.6 s1.33x
queries sent2,0002,000
cpu643.8 s1,624.8 s2.52x
cpu, % of wall420%798%
waiting for a core102.3 ms13,447.9 ms131.48x
migrations06,409
faults min/maj0 / 01,478 / 0
peak RSS29.3 GiB43.6 GiB1.49x
disk written00
disk read00
1.33x is 2.62x less work per query x 0.53x cores busy during the row (4.20 against 7.98) x 1.00x clock x 0.97x counted on-CPU share: the middle term is occupancy, not search speed brute force over the whole collection, no index involved

W10-ef32 recall control, ef=32 (latency only) 1.44x

ef=32 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second10,6257,3581.44x
recall@100.89120.8945~equal
client p50737 µs1.00 ms1.36x
client p991.20 ms1.98 ms1.65x
client p99.91.67 ms2.54 ms1.52x
server p50641 µs816 µs1.27x
server p991.11 ms1.72 ms1.56x
wall clock4.7 s6.8 s1.44x
queries sent50,00050,000
cpu32.5 s40.4 s1.24x
cpu, % of wall685%591%
waiting for a core3.8 ms15,065.5 ms3,983.78x
migrations084,121
faults min/maj0 / 01,261 / 0
peak RSS29.4 GiB43.6 GiB1.49x
disk written00
disk read00

W10-ef64 recall control, ef=64 (latency only) 1.33x

ef=64 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second6,4814,8651.33x
recall@100.93990.9441~equal
client p501.23 ms1.54 ms1.25x
client p991.92 ms3.02 ms1.57x
client p99.92.39 ms3.73 ms1.56x
server p501.13 ms1.32 ms1.17x
server p991.81 ms2.66 ms1.47x
wall clock7.8 s10.3 s1.33x
queries sent50,00050,000
cpu54.5 s65.3 s1.20x
cpu, % of wall701%633%
waiting for a core5.4 ms38,874.7 ms7,258.48x
migrations0155,076
faults min/maj0 / 01,194 / 0
peak RSS29.4 GiB43.6 GiB1.49x
disk written00
disk read00

W10-ef128 recall control, ef=128 (latency only) 1.24x

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second3,7993,0741.24x
recall@100.96670.9694~equal
client p502.11 ms2.52 ms1.19x
client p993.27 ms4.63 ms1.41x
client p99.93.76 ms5.69 ms1.51x
server p502.00 ms2.26 ms1.13x
server p993.16 ms4.22 ms1.34x
wall clock13.2 s16.3 s1.24x
queries sent50,00050,000
cpu93.4 s107.3 s1.15x
cpu, % of wall707%659%
waiting for a core8.8 ms58,175.7 ms6,645.12x
migrations0178,808
faults min/maj0 / 0647 / 0
peak RSS29.4 GiB43.6 GiB1.49x
disk written00
disk read00

W10-ef256 recall control, ef=256 (latency only) 1.18x

ef=256 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,1811,8541.18x
recall@100.98120.9822~equal
client p503.66 ms4.23 ms1.16x
client p995.78 ms7.42 ms1.28x
client p99.96.60 ms8.94 ms1.35x
server p503.54 ms3.94 ms1.11x
server p995.67 ms6.98 ms1.23x
wall clock23.0 s27.0 s1.18x
queries sent50,00050,000
cpu162.2 s186.5 s1.15x
cpu, % of wall706%690%
waiting for a core15.2 ms81,122.1 ms5,330.22x
migrations0193,845
faults min/maj0 / 0664 / 0
peak RSS29.4 GiB43.6 GiB1.49x
disk written00
disk read00

W10-ef512 recall control, ef=512 (latency only) 1.16x

ef=512 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,2351,0681.16x
recall@100.98920.9901~equal
client p506.44 ms7.39 ms1.15x
client p9910.35 ms12.24 ms1.18x
client p99.911.85 ms14.06 ms1.19x
server p506.32 ms7.01 ms1.11x
server p9910.22 ms11.74 ms1.15x
wall clock40.5 s46.9 s1.16x
queries sent50,00050,000
cpu285.5 s335.4 s1.17x
cpu, % of wall704%716%
waiting for a core25.5 ms110,818.9 ms4,337.95x
migrations0195,531
faults min/maj3 / 01,846 / 0
peak RSS29.4 GiB43.6 GiB1.49x
disk written00
disk read00

W12-upload filtered search: load 200,000 with payloads

upload · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=200,000
metricstrawmANNQdrant
wall clock45.8 s98.6 s2.15x
of which upload1.5 s5.8 s3.79x
of which index wait39.1 s90.2 s2.31x
time to Green39.1 s90.2 s2.31x
cpu297.7 s520.5 s1.75x
cpu, % of wall650%528%
waiting for a core166.6 ms10,982.4 ms65.91x
migrations07,863
faults min/maj445,699 / 34372,899 / 446,815
peak RSS31.6 GiB43.6 GiB1.38x
disk written1.1 GiB5.3 GiB
disk read588.0 KiB61.1 MiB

W12-sel1 filtered search, one keyword (~1% of bench12) 1.16x

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,8691,6161.16x
recall@101.00000.9995~equal
client p501.06 ms1.24 ms1.16x
client p991.13 ms1.44 ms1.27x
client p99.91.19 ms1.51 ms1.28x
server p50967 µs1.12 ms1.15x
server p991.02 ms1.30 ms1.27x
wall clock26.8 s31.0 s1.16x
queries sent50,00050,000
cpu49.5 s62.7 s1.27x
cpu, % of wall185%202%
waiting for a core4.5 ms21.6 ms4.76x
migrations017,479
faults min/maj0 / 01,065 / 0
peak RSS31.6 GiB43.6 GiB1.38x
disk written00
disk read00
filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set

W12-sel10 filtered search, any of 10 keywords (~10% of bench12)

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second822823-
recall@100.98930.9720
client p502.49 ms2.44 ms
client p993.03 ms2.78 ms
client p99.93.21 ms2.94 ms
server p502.38 ms2.29 ms
server p992.93 ms2.62 ms
wall clock60.9 s60.8 s
queries sent50,00050,000
cpu118.1 s122.3 s
cpu, % of wall194%201%
waiting for a core6.9 ms131.2 ms
migrations019,113
faults min/maj0 / 0997 / 0
peak RSS31.6 GiB43.6 GiB
disk written00
disk read00
recall unequal: 0.9893 vs 0.9720; §7.4 compares at equal recall filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set

W12-sel1-ef32 filtered recall control, one keyword, ef=32 (latency only)

ef=32 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,8643,397-
recall@101.00000.9804
client p501.07 ms585 µs
client p991.13 ms687 µs
client p99.91.18 ms752 µs
server p50969 µs473 µs
server p991.02 ms561 µs
wall clock26.9 s14.8 s
queries sent50,00050,000
cpu49.5 s29.8 s
cpu, % of wall184%202%
waiting for a core1.1 ms11.6 ms
migrations010,549
faults min/maj0 / 0304 / 0
peak RSS31.6 GiB43.6 GiB
disk written00
disk read00
recall unequal: 1.0000 vs 0.9804; §7.4 compares at equal recall

W12-sel1-ef64 filtered recall control, one keyword, ef=64 (latency only) 0.77x

ef=64 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,8622,4120.77x
recall@101.00000.9965~equal
client p501.07 ms826 µs0.77x
client p991.13 ms956 µs0.85x
client p99.91.19 ms1.02 ms0.86x
server p50969 µs712 µs0.73x
server p991.02 ms829 µs0.81x
wall clock26.9 s20.8 s0.77x
queries sent50,00050,000
cpu49.4 s42.0 s0.85x
cpu, % of wall184%202%
waiting for a core1.7 ms10.9 ms6.25x
migrations013,463
faults min/maj0 / 0309 / 0
peak RSS31.6 GiB43.6 GiB1.38x
disk written00
disk read00

W12-sel1-ef128 filtered recall control, one keyword, ef=128 (latency only) 1.14x

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,8641,6361.14x
recall@101.00000.9995~equal
client p501.07 ms1.22 ms1.14x
client p991.13 ms1.41 ms1.25x
client p99.91.17 ms1.49 ms1.28x
server p50969 µs1.10 ms1.14x
server p991.02 ms1.28 ms1.26x
wall clock26.9 s30.6 s1.14x
queries sent50,00050,000
cpu49.5 s61.8 s1.25x
cpu, % of wall184%202%
waiting for a core1.4 ms19.6 ms13.77x
migrations017,422
faults min/maj0 / 0581 / 0
peak RSS31.6 GiB43.6 GiB1.38x
disk written00
disk read00

W12-sel1-ef256 filtered recall control, one keyword, ef=256 (latency only) 1.72x

ef=256 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,8641,0831.72x
recall@101.00000.9999~equal
client p501.07 ms1.85 ms1.73x
client p991.13 ms2.09 ms1.85x
client p99.91.17 ms2.19 ms1.88x
server p50968 µs1.71 ms1.77x
server p991.02 ms1.94 ms1.90x
wall clock26.9 s46.2 s1.72x
queries sent50,00050,000
cpu49.5 s93.2 s1.88x
cpu, % of wall184%202%
waiting for a core1.1 ms118.2 ms108.00x
migrations018,434
faults min/maj0 / 0723 / 0
peak RSS31.6 GiB43.6 GiB1.38x
disk written00
disk read00

W12-sel1-ef512 filtered recall control, one keyword, ef=512 (latency only) 2.45x

ef=512 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,8617582.45x
recall@101.00001.0000~equal
client p501.07 ms2.63 ms2.46x
client p991.13 ms2.88 ms2.54x
client p99.91.17 ms3.17 ms2.70x
server p50971 µs2.46 ms2.54x
server p991.02 ms2.68 ms2.63x
wall clock26.9 s66.0 s2.45x
queries sent50,00050,000
cpu49.4 s131.4 s2.66x
cpu, % of wall184%199%
waiting for a core1.9 ms226.9 ms118.69x
migrations016,525
faults min/maj0 / 01,210 / 0
peak RSS31.6 GiB43.6 GiB1.38x
disk written00
disk read00

W12-sel10-ef32 filtered recall control, any of 10 keywords, ef=32 (latency only)

ef=32 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,9742,259-
recall@100.93830.7668
client p501.02 ms872 µs
client p991.32 ms1.14 ms
client p99.91.56 ms1.30 ms
server p50923 µs753 µs
server p991.22 ms1.02 ms
wall clock25.4 s22.2 s
queries sent50,00050,000
cpu47.0 s44.9 s
cpu, % of wall185%202%
waiting for a core7.4 ms10.7 ms
migrations015,993
faults min/maj0 / 0652 / 0
peak RSS31.6 GiB43.6 GiB
disk written00
disk read00
recall unequal: 0.9383 vs 0.7668; §7.4 compares at equal recall

W12-sel10-ef64 filtered recall control, any of 10 keywords, ef=64 (latency only)

ef=64 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,3121,400-
recall@100.97280.9005
client p501.55 ms1.42 ms
client p991.90 ms1.73 ms
client p99.92.16 ms1.89 ms
server p501.45 ms1.29 ms
server p991.80 ms1.58 ms
wall clock38.1 s35.8 s
queries sent50,00050,000
cpu72.6 s72.3 s
cpu, % of wall190%202%
waiting for a core10.3 ms38.0 ms
migrations019,122
faults min/maj0 / 0736 / 0
peak RSS31.6 GiB43.6 GiB
disk written00
disk read00
recall unequal: 0.9728 vs 0.9005; §7.4 compares at equal recall

W12-sel10-ef128 filtered recall control, any of 10 keywords, ef=128 (latency only)

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second822822-
recall@100.98930.9720
client p502.49 ms2.44 ms
client p993.03 ms2.79 ms
client p99.93.22 ms2.93 ms
server p502.38 ms2.29 ms
server p992.93 ms2.63 ms
wall clock60.9 s60.9 s
queries sent50,00050,000
cpu118.0 s122.5 s
cpu, % of wall194%201%
waiting for a core10.3 ms154.7 ms
migrations019,663
faults min/maj0 / 01,051 / 0
peak RSS31.6 GiB43.6 GiB
disk written00
disk read00
recall unequal: 0.9893 vs 0.9720; §7.4 compares at equal recall

W12-sel10-ef256 filtered recall control, any of 10 keywords, ef=256 (latency only) parity

ef=256 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second480486parity
recall@100.99590.9945~equal
client p504.23 ms4.13 ms0.98x
client p995.27 ms4.75 ms0.90x
client p99.95.55 ms5.11 ms0.92x
server p504.03 ms3.94 ms0.98x
server p995.03 ms4.52 ms0.90x
wall clock104.3 s103.0 s~equal
queries sent50,00050,000
cpu201.5 s205.8 s1.02x
cpu, % of wall193%200%
waiting for a core17.9 ms163.6 ms9.15x
migrations022,176
faults min/maj0 / 02,134 / 0
peak RSS31.6 GiB43.6 GiB1.38x
disk written00
disk read00
within the ±4.5% band this dataset's noise floor puts on W12-sel10-ef256: no measured difference, not a small one

W12-sel10-ef512 filtered recall control, any of 10 keywords, ef=512 (latency only) parity

ef=512 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second294286parity
recall@101.00000.9991~equal
client p506.81 ms6.99 ms1.03x
client p997.56 ms8.12 ms1.07x
client p99.97.85 ms8.46 ms1.08x
server p506.42 ms6.71 ms1.05x
server p996.98 ms7.82 ms1.12x
wall clock170.2 s174.8 s1.03x
queries sent50,00050,000
cpu325.0 s348.9 s1.07x
cpu, % of wall191%200%
waiting for a core38.1 ms140.4 ms3.69x
migrations023,563
faults min/maj0 / 04,057 / 0
peak RSS31.6 GiB43.6 GiB1.38x
disk written00
disk read00
within the ±4.9% band this dataset's noise floor puts on W12-sel10-ef512: no measured difference, not a small one

W13 scroll / pagination 1.09x

closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=200,000
metricstrawmANNQdrant
queries/second51,14146,8521.09x
client p50137 µs157 µs1.15x
client p99230 µs293 µs1.27x
client p99.9267 µs446 µs1.67x
server p5011 µs28 µs2.55x
server p9916 µs90 µs5.60x
wall clock4.0 s2.2 s0.55x
queries sent200,000100,000
cpu5.6 s9.1 s1.63x
cpu, % of wall138%413%
waiting for a core4.2 ms853.7 ms201.15x
migrations0108,063
faults min/maj0 / 0780 / 0
peak RSS31.6 GiB43.6 GiB1.38x
disk written00
disk read00
strawmANN: -n 200000 (4x the table's 50000) Qdrant: -n 100000 (2x the table's 50000) strawmANN: 0.13 ms of the client's 0.14 ms p50 is not server time (12x), so 92% of this row is the load generator and the socket Qdrant: 0.13 ms of the client's 0.16 ms p50 is not server time (6x), so 82% of this row is the load generator and the socket 1.09x is 3.25x less work per query x 0.33x cores busy during the row (1.40 against 4.23) x 1.00x clock x 1.02x counted on-CPU share: the middle term is occupancy, not search speed

W11-steady mixed read/write below the rebuild threshold: search bench2 while 49,500 synthetic points append

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=31,122
metricstrawmANNQdrant
queries/second580504-
client p509.02 ms31.01 ms
client p9934.05 ms68.87 ms
client p99.9114.14 ms82.97 ms
server p508.39 ms30.23 ms
server p9933.20 ms67.88 ms
wall clock45.0 s114.4 s
queries sent31,12231,122
cpu351.5 s900.4 s
cpu, % of wall781%787%
waiting for a core263,490.0 ms784,454.5 ms
migrations1,11858,423
faults min/maj580,619 / 91,221889,903 / 147,520
peak RSS35.1 GiB43.6 GiB
disk written290.0 MiB12.5 GiB
disk read71.8 MiB161.2 MiB
strawmANN: qps over the 26.1 s the append ran (726 over the whole search); append 1,900 points/s Qdrant: qps over the 26.0 s the append ran (272 over the whole search); append 1,900 points/s search-during-write; no recall join

W11 mixed read/write: search bench2 while 198,000 synthetic points append (runs last)

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=7,945
metricstrawmANNQdrant
queries/second25199-
client p5021.89 ms70.85 ms
client p99113.84 ms186.56 ms
client p99.9140.08 ms1,871.42 ms
server p5020.74 ms69.57 ms
server p99112.83 ms184.99 ms
wall clock31.7 s80.2 s
queries sent7,9457,945
cpu476.4 s652.0 s
cpu, % of wall1,503%813%
waiting for a core208,739.9 ms518,000.9 ms
migrations80625,478
faults min/maj1,020,754 / 298,410769,832 / 324,970
peak RSS35.1 GiB43.6 GiB
disk written1.1 GiB8.9 GiB
disk read13.4 MiB77.2 MiB
strawmANN: qps over the 31.7 s the append ran (251 over the whole search); write overlap 53%; search covered 53% of the append; append 3,300 points/s Qdrant: qps over the 82.1 s the append ran (97 over the whole search); search covered 97% of the append; append 2,396 points/s search-during-write; no recall join; the search saw under 90% of strawmANN's append, so its qps is the append's start and not all of it
Host discipline the §7.1 gate, per check, and ambient load per row

Ambient load per row

The §7.1 gate checks the machine once, at the start. Colour is the engine, as everywhere else; a hatched bar with a red edge is a row that had another process on the box while it ran, which the gate cannot see.

strawmANN gate pass

environment hash 373e8531fd3d1bf6
okAMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical)
okgovernor=performance
okboost disabled
oksmt=on
declared rather than required (§7.1 asks for SMT on or off, not
for off), and hashed, so these rows never share a chart with
SMT-off ones
cpu N and cpu N+12 share a core (12 pairs);
a --server-cpus or --client-cpus set holding both cpus of a
pair measures the engine on half the cores it names
okprofile=as-deployed (isolcpus=none nohz_full=none)
the scheduler is running as it ships, so these numbers describe a
deployment rather than the engine in isolation; they may not be
compared against an `isolated` run, and the environment hash
enforces that
okthp=madvise numa_nodes=1 numa_balancing=0
okperf_event_paranoid=-1
oktask_delayacct=1, so time blocked on a device is measurable
okperf at /usr/bin/perf, so `workloads.py run --perf` can attach
okcgroup io controller reaches the engines' own scope, so block-layer read/write operations are measurable
okzig=0.16.0
busy: Xorg(5%)
measurement profile: as-deployed

Qdrant gate pass

environment hash 373e8531fd3d1bf6
okAMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical)
okgovernor=performance
okboost disabled
oksmt=on
declared rather than required (§7.1 asks for SMT on or off, not
for off), and hashed, so these rows never share a chart with
SMT-off ones
cpu N and cpu N+12 share a core (12 pairs);
a --server-cpus or --client-cpus set holding both cpus of a
pair measures the engine on half the cores it names
okprofile=as-deployed (isolcpus=none nohz_full=none)
the scheduler is running as it ships, so these numbers describe a
deployment rather than the engine in isolation; they may not be
compared against an `isolated` run, and the environment hash
enforces that
okthp=madvise numa_nodes=1 numa_balancing=0
okperf_event_paranoid=-1
oktask_delayacct=1, so time blocked on a device is measurable
okperf at /usr/bin/perf, so `workloads.py run --perf` can attach
okcgroup io controller reaches the engines' own scope, so block-layer read/write operations are measurable
okzig=0.16.0
busy: Xorg(6%)
measurement profile: as-deployed
What this does not establish

No Qdrant developer has used this tool. Nothing here has been run, checked against something already known, or disagreed with by anyone outside the project: every number has one author and one reviewer, and they are the same person. A result you can refute is more useful to us than one you accept.

It has never been pointed at a real Qdrant regression. The harness detects a known ISA slowdown and correctly reports no difference on a control row. That is internal consistency, not evidence it would flag a regression in your tree or stay quiet through a refactor. bench/harness/qdrant_ab.py exists to run two of your commits through it blind, and that experiment has not been done.

Part of the measured throughput gap is kernel width, not architecture. Read from a dev checkout: Qdrant's distance path tops out at four 256-bit accumulators and has no AVX-512, while these kernels use 512-bit ones on a machine that has them. That is a real difference and it is not the same claim as "the design is faster".

The engines are not equivalent, by construction. §1 removes sharding, replication, consensus, snapshots, sparse vectors, multivectors and disk-resident operation — permanently. Payload storage and filtering were in that list until 2026-09-03 and are now built (M7), so W12 measures a filtered search with a keyword index on both engines. Anything still unbuilt answers UNIMPLEMENTED naming the construct and reports n/a, never a silent degradation. A strawman that was not faster would mean it was badly built; the question is by how much, and where the model was wrong.

Reproducing this

Everything below runs from a clone. The engine has no dependencies; the analysis path is a uv project pinned by bench/uv.lock.

scripts/doctor.py                    # what can this machine measure?
bench/setup.py check                 # §7.1 host gate: governor, boost, isolation, idle
conformance/datasets/datasets.py fetch

# one engine at a time, §7.1
zig build -Doptimize=ReleaseFast
# --connections matters: W4 opens ~32 sockets (-t 16 -c 2) and a server with
# fewer closes the excess before the HTTP/2 preface, which the client reports
# only as "transport error".
./zig-out/bin/strawmann --port 6334 --connections 64
bench/harness/workloads.py run http://localhost:6334 strawmann --sink --report

# the other half of every number: recall, and the licence to compare at all
cd conformance
cargo run --release -- relevance --engine http://localhost:6334 --label strawmann \
  --base $DATA/sift1m.fbin --queries $DATA/sift1m_query.fbin \
  --ground-truth $DATA/gt/sift1m.euclid.k100.gt.json --metric euclid \
  --ef 32 64 128 256 512 --json ../bench/results/strawmann/recall.json
cargo run --release -- differ --strawmann http://localhost:6334 --qdrant http://localhost:6434 \
  --base $DATA/sift1m.fbin --queries $DATA/sift1m_query.fbin \
  --ground-truth $DATA/gt/sift1m.euclid.k100.gt.json \
  --json ../bench/results/strawmann/conformance.json

uv run --project bench bench/harness/report.py strawmann qdrant

Disagreements are the point. docs/spec.md is what all of this is measured against, docs/bugs.md lists the measurement bugs found so far (each produced a plausible wrong number rather than a failure), and docs/validation.md records what has and has not been checked.