strawmANN benchmark report

strawmANN against Qdrant · AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical) · generated 2026-09-29 10:29:45 CEST

Dataset
h-and-m-2048-angular-filters
Vectors
105,100 × 2,048 · 10,000 held-out queries
Metric
cosine
Storage datatype
strawmANN float32 · Qdrant float32 (default)
Quantization
none (fp32) · separate rows measure binary (1 bit/dim), PQ (product), SQ8 scalar
strawmANN: 43/43 ok Qdrant: 43/43 ok

Summary

At equal recall, strawmANN serves 1.30x to 1.66x Qdrant's throughput. That range spans the recall levels measured; allowing for the uncertainty in each recall figure widens it to 0.72x to 2.61x.

throughput at equal recall
1.30 to 1.66x
strawmANN over Qdrant, recall@10 0.994 to 0.999
p99 latency at 90% load
strawmANN2.57 ms
Qdrant8.26 ms
W4-sat90, open loop, each engine's own saturation
upload and index build
strawmANN24 s
Qdrant26 s
W1 + W2
peak memory (RSS)
strawmANN5.6 GiB
Qdrant8.6 GiB
largest before the concurrent-write rows (W11); of it, anonymous 609.2 MiB / 550.4 MiB
written to disk
strawmANN4.8 GiB
Qdrant22.3 GiB
before the concurrent-write rows (W11); they wrote 205.8 MiB / 31.6 GiB more

Where strawmANN is slower, past the noise floor:

Throughput against recall

Read it vertically: at any recall both engines reach, the higher curve is faster. Throughput from bfb W10, recall from the conformance sweep, joined on ef. ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.

Every figure is the median of 3 passes per engine, run alternately (A/B/A/B). What this does not establish.

Glossary
recall@10
Of the ten nearest neighbours a query really has, the share the engine returned. 1.0 is a perfect answer. "Really has" is settled by an exhaustive fp64 search, not by the other engine.
ef
The candidate-list size: how many nodes the index keeps in play while searching. Larger is slower and more accurate. It does not mean the same amount of work in both engines, which is why the headline holds recall fixed instead.
MRDE
Mean relative distance error: when the engine returns a wrong neighbour, how much further away it is than the right one. Small numbers mean the misses were near-misses.
qps
Queries per second. On a batched row one request carries several queries, and the request rate is shown under it.
p50 / p99 / p99.9
The latency half of requests beat, that 99% beat, that 99.9% beat. The tail is what a user notices.
closed / open loop
A closed loop sends the next request only when the last one comes back, so a slow server receives less work and its tail looks better than it is. An open loop sends at a fixed rate regardless.
noise floor
How much a number moves between identical runs on this machine. A ratio inside it is shown grey with ≈: no measured difference. Hover a ratio to see its band.
conformance tier
What the two engines were shown to agree on before any speed was quoted. T1 licenses a single engine's own numbers; T3 licenses comparing the two, and requires their recall to be statistically indistinguishable.
§ numbers
Sections of docs/spec.md, the written rule each claim in the appendix is measured against. §7.1 is the host gate, §7.4 the comparison rules, §8 what may be published.

Throughput

Queries per second; higher is better. The ratio is strawmANN over Qdrant: green is faster, red slower, grey ≈ inside the noise floor (hover a ratio for its band). A dash means the pair is not compared, and the note says why. An ef sweep is one row showing its range; every measurement is under All throughput rows.

workloadstrawmANNQdrantrationotes
W3search, fp32, single query1,2819861.30x
W4search, saturating (closed loop)4,6565,6830.82x
W5search batched (16 distinct dataset queries per request)4,656291 requests/s2,164136 requests/s2.15x
W6quantized: scalar3,6212,7261.33x
W6 ef 32 to 512SQ8 recall control (latency only)1,423 to 7,6191,005 to 5,2681.27 to 1.45x
W7quantized: binary + oversampling4,3524,354parity
W8quantized: PQ1,4521,1211.30x
W9exact / brute force7171parity
W10 ef 32 to 512recall control (latency only)1,731 to 11,9671,959 to 10,4280.88 to 1.15x
W12-sel1filtered search, one keyword (~1% of bench12)2,5942,503parity
W12-sel10filtered search, any of 10 keywords (~10% of bench12)1,1471,280parity
W12-sel1 ef 32 to 512filtered recall control, one keyword (latency only)2,593 to 2,6121,137 to 4,5990.57 to 2.29x1 of 5 points not compared
W12-sel10 ef 32 to 512filtered recall control, any of 10 keywords (latency only)408 to 2,583486 to 2,9980.52 to 0.85x3 of 5 points not compared recall differs
W13scroll / pagination50,84546,4501.09xstrawmANN: 92% is client and socket, not server Qdrant: 82% is client and socket, not server
W11-steadymixed read/write below the rebuild threshold: search bench2 while 5,255 synthetic points append2,3883,857-search-during-write; no recall join
W11mixed read/write: search bench2 while 21,020 synthetic points append (runs last)6233,453-search-during-write; no recall join

Throughput by workload

Higher is better. Hatched bars are rows the table does not compare (W7, W9, W11): each bar is that engine's own rate, and the pair is not a result.

Recall

recall@10 against an exact fp64 search, not against the other engine. Equal ef is not equal work in the two engines, so the comparison that counts holds recall fixed and compares throughput there.

At matched recall

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.9943strawmANN at ef=32, Qdrant interpolated11,9677,1931.66x1.50 – 1.85x
0.9945Qdrant at ef=64, strawmANN interpolated11,7137,1411.64x1.21 – 2.22x
0.9970Qdrant at ef=128, strawmANN interpolated7,7404,7911.62x1.33 – 1.96x
0.9972strawmANN at ef=64, Qdrant interpolated7,4764,5191.65x1.20 – 2.27x
0.9985Qdrant at ef=256, strawmANN interpolated4,7333,1201.52x1.12 – 2.05x
0.9986strawmANN at ef=128, Qdrant interpolated4,6363,0061.54x0.91 – 2.61x
0.9992strawmANN at ef=256, Qdrant interpolated2,8522,1101.35x0.87 – 2.11x
0.9993Qdrant at ef=512, strawmANN interpolated2,5411,9591.30x0.72 – 2.35x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W10 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column is why the ratio is not a result on its own. The ratio treats the anchor's recall as exact; it is an estimate, and moving it across its 95% interval moves the interpolated rate with it — near the top of the sweep 0.003 of recall spans a factor of 1.8, so two decimals there quote the interpolation. Even that is the narrow reading: it moves one recall and not the other, adjacent anchors share bracketing segments, and the throughputs behind it are medians of the run's passes.

At matched recall, SQ8

The same reading over the scalar-quantized collection. Under the `pool` policy both engines rescore `ef` candidates, so W6 is compared row by row as well; held at equal recall, the reading does not depend on the rows landing in one recall band. Note where each engine's curve stops, because a recall only one of them reaches is the more useful fact about an encoding than any ratio.

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.9947strawmANN at ef=32, Qdrant interpolated7,6194,0031.90x1.74 – 2.09x
0.9948Qdrant at ef=64, strawmANN interpolated7,5093,9851.88x1.42 – 2.49x
0.9974Qdrant at ef=128, strawmANN interpolated5,1772,7311.90x1.58 – 2.28x
0.9975strawmANN at ef=64, Qdrant interpolated5,0582,5721.97x1.14 – 3.39x
0.9986Qdrant at ef=256, strawmANN interpolated3,9101,7342.26x1.81 – 2.81x
0.9989strawmANN at ef=128, Qdrant interpolated3,6151,3652.65x1.18 – 5.94x
0.9993Qdrant at ef=512, strawmANN interpolated2,7301,0052.72x1.62 – 4.54x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W6 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column reads as it does in the fp32 table above.

At matched recall, filtered to 10%

The same reading under a keyword filter over the points the condition matched, 10,608 in strawmANN and 10,565 in Qdrant (bfb draws the keyword payloads unseeded at each upload). The per-row table refuses W12-sel10 a ratio because the two engines land just outside the recall band at the one ef it measures; held at equal recall instead, the comparison exists at every recall both engines reach. Note the shape rather than any single number: one engine's curve is flat in ef and the other's is steep, so where you match decides the ratio.

recall@10measured atstrawmANN q/sQdrant q/sstrawmANN / Qdrantrange over recall CI
0.9929strawmANN at ef=32, Qdrant interpolated2,5831,3201.96x1.87 – 2.04x
0.9952Qdrant at ef=128, strawmANN interpolated2,1381,2681.69x1.45 – 1.96x
0.9975strawmANN at ef=64, Qdrant interpolated1,7619591.84x1.56 – 2.17x
0.9991strawmANN at ef=128, Qdrant interpolated1,1467971.44x1.28 – 1.61x
0.9992Qdrant at ef=256, strawmANN interpolated9927841.26x0.59 – 2.71x
0.9999Qdrant at ef=512, strawmANN interpolated4824860.99x0.65 – 1.50x
How this is computed

Interpolated linear in log(q/s) between the two bracketing measurements, nothing extrapolated past the measured range. Throughput is bfb's W12-sel10 sweep, recall the conformance sweep, joined on ef — same m and ef_construct, but separate builds, so the pairing assumes two builds with identical parameters are equivalent.

The last column reads as it does in the fp32 table above.

Recall by ef

efstrawmANNQdrant
recall@1recall@10recall@100MRDErecall@1recall@10recall@100MRDE
ef=320.99780.9943-1.25e-030.99560.9877-1.75e-03
ef=640.99900.9972-6.72e-040.99800.9945-7.16e-04
ef=1280.99960.99860.99233.47e-040.99920.99700.98664.52e-04
ef=2560.99960.99920.99753.14e-040.99950.99850.99591.62e-04
ef=5120.99980.99970.99934.57e-050.99960.99930.99841.36e-04
How recall was measured

10,000 held-out queries, limit 10, ε=6.556510925292969e-07, against the fp64 oracle rather than against the other engine. Base checksum 5a7faba68c84ce7c: the same corpus the latency rows were measured on. Each pass rebuilds the collection. Across 3 independent builds of it, recall@10 at ef=512 spread by: strawmANN 0.00016; Qdrant 0.00012. Over the same builds the engine reported unreachable nodes: strawmANN 0 of 127,843 — a node nothing points at is invisible at any ef, so that is where a graph-quality difference shows rather than being inferred from the recall beside it. Level seed: strawmANN 0x57ea3111. Recorded as provenance: under this engine's level draw, six builds at four seeds sit within 0.00003 of recall@10 at ef 512, well inside the spread above.

What each quantization costs in recall

Colour is the encoding here, not the engine — the engines are the line style. Each sweep is joined on the quantization parameters it actually sent, so a curve speaks only for the search its row ran. Higher is better, and the encodings are not free: read this against the throughput their rows bought.

Latency

Client-side round trip. A closed loop understates the tail (a stalled server stops receiving requests), so this keeps the fixed-rate, open-loop rows, which offer the same load to both engines, plus single-client W3 and batched W5. Every row is under All latency rows.

workloadstrawmANNQdrant
p50p99p99.9p50p99p99.9
W3 search, fp32, single query759 µs1.40 ms1.73 ms1.00 ms1.43 ms1.54 ms
W4-sat50 search, fixed rate at 50% of saturation (open loop)889 µs1.57 ms1.80 ms1.12 ms2.09 ms2.49 ms
W4-sat70 search, fixed rate at 70% of saturation (open loop)1.01 ms2.12 ms2.47 ms1.30 ms2.65 ms3.56 ms
W4-sat90 search, fixed rate at 90% of saturation (open loop)1.15 ms2.57 ms3.51 ms1.91 ms8.26 ms13.36 ms
W5 search batched (16 distinct dataset queries per request)6.86 ms8.12 ms8.55 ms14.78 ms16.64 ms17.30 ms

Ingest and index build

Qdrant indexes while it ingests, so its upload time already contains most of the indexing; strawmANN uploads raw and builds afterwards. Compare the sum, not the upload line.

Where the load time goes

Solid is upload, hatched is the wait for Green. The bar's whole length is the sum this section asks you to compare; the split is why the upload line alone is not comparable between these engines.

workloadstrawmANNQdrant
W0-uploadd=4 floor: graph traversal with the distance taken out4.01 s4.01 s †
W1ingest throughput (no index wait)1.04 s2.57 s
W2index build, time to Green23.08 s23.06 s
W6-uploadsearch, SQ8 scalar quantization85.20 s19.05 s
W7-uploadsearch, binary quantization25.08 s11.03 s
W8-uploadsearch, product quantization173.56 s145.37 s
W12-uploadupload23.07 s44.11 s
W1 + W2upload and index, together24.1 s25.6 s1.06x

† at the load generator's polling floor. bfb decides a collection is Green by polling once a second and requiring three consecutive Green replies, having slept a second before the first poll, so no Time-to-Green it can report is below 3 s regardless of how fast the build was. A marked cell is an upper bound on the build and a measurement of the polling loop. The bias is a constant added to both engines, so the difference between them survives it and the ratio does not — and it is the faster engine the ratio understates.

Memory and disk

What each engine held in memory and moved to and from disk over the whole run. Per-row figures are under Storage and I/O.

strawmANNQdrant
storage on disk6.0 GiB4.6 GiB
peak memory (RSS)5.6 GiB8.6 GiB
read from disk20.0 KiB282.4 MiB
written to disk4.8 GiB22.3 GiB

Peak memory and storage on disk are read before the concurrent-write rows (W11): an engine rewriting segments maps old and new files at once, and RSS counts each mapping. Bytes written are summed over the same rows, since W11's volume is set by the harness's write rate; the other disk figures are totals over every row.

Appendix

How the run was set up, every row of every table, and the diagnostics behind them. Charts on a logarithmic axis say so on the axis: read the positions there, not the distances.

Run conditions and conformance

Measured as-deployed: the scheduler was left as it ships, so a difference is what a user would see rather than the engine in isolation, and ambient load is part of the measurement. Qdrant ran equal-work: asked for one populated graph (--segments 1), as strawmANN serves, so ef means the same thing on both sides. This is the configuration §8's comparative licensing is built around, and it is a control rather than a deployment. Its optimizer treats the count as a target; what it actually held is read back from the engine and stated under Collection config.

Conformance tier reached: T4 quantization fidelity. T1 passed, so single-engine performance rows are licensed (§8). T3 passed, so a strawmANN-vs-Qdrant throughput comparison is licensed: the engines are at equal recall (§7.4). Conformance hash b8c6511753f0e22d.

On exact search the two engines' scores differ by at most 2.980e-07 (p99 1.192e-07), against the §8.4 floor absolute ε of 6.557e-07 — the expected summation error for this metric and dimension, used because no calibrated cell exists for them; the conformance calibrate command derives one from measured cross-ISA spread. §8.1: bit-exactness is unachievable between two different summation orders, so this — equal values within a measured tolerance — is the claim the speed rests on.

The conformance harness's single-client rate agrees on the ordering (2.01x to 2.05x); a different instrument, so no ratio is formed from it.

Every figure is the median of 3 passes per row per engine, alternated A/B/A/B (§7.4), and the spread is this run's own over 36 of 43 rows. The ingest and index-build rows (W0-upload, W1, W2, W6-upload, W7-upload, W8-upload, W12-upload) have none, so no build time carries a verdict. Each pass rebuilds the index, so the bands include build variance and are wider than a floor measured against one standing graph.

Units are queries per second. bfb reports rps, which counts batch requests: at --search-batch-size 16 the two differ by 16x, and reading one as the other once turned a 2.2x speedup into an apparent 7x regression. The tables show the request rate wherever it diverges.

The open-loop rows are not a speed. W4-sat50/70/90 offer a fixed fraction of measured saturation; serving it means the engine kept up, not that it was faster, so no ratio is printed. They exist because --parallel is a closed loop, where a stalled server stops receiving requests and understates its own tail.

ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.

What was measured builds, dataset, host
strawmANN
sm-hnm-perf-0929
Qdrant
qd-hnm-perf-0929
enginecommit e14920833a08
ReleaseFast, native build, vnni on
binary sha256 eeb45cddf8af63e5
built 2026-09-28T12:30:05Z with zig 0.16.0
version 1.19.2-dev
binary ~/.cache/strawmann/qdrant-dbeb0f73/qdrant
sha256 dbeb0f73dea2d371 build 878843e6 (from the server's banner, not a checkout)
native binary outside a checkout: sha256 is the identity (§8.9)
network native no container in the path, like strawmann
measured2026-09-29T04:30:25Z to 2026-09-29T07:20:31Z (3 passes)2026-09-29T05:05:58Z to 2026-09-29T07:55:51Z (3 passes)
profileas-deployed · both engines
load generatorbfb dev @ fc6632e5 (qdrant/bfb#176; carries #172, findings 32's --rps reaping fix) · both engines
clientqdrant-client 1.16.1-dev (git dev branch) · both engines
launched as~/Workspace/strawmann/zig-out/bin/strawmann --port 6334 --capacity 131375 --connections 64 --workers 7 --io-threads 1 --pin --cpus 4-11 --data-dir ~/.cache/strawmann/strawmann-storage --default-placement cached~/.cache/strawmann/qdrant-dbeb0f73/qdrant
cores the engine could use4-11 (8 cores) requested by the harness
4-11 (8 cores) observed on 8 threads
0-23 (24 cores) observed on 1 thread
4-11 (8 cores) requested by the harness
4-11 (8 cores) observed on 50 threads

Dataset

h-and-m-2048-angular-filters, 105100 × 2048, cosine, 10000 held-out queries
ground truth: shipped, and diffed against our fp64 recompute
1 file, each pinned by sha256 in datasets.json

Host §7.1 gate pass

AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M
24 logical cores · 58.6 GB · kernel 7.0.0-34-generic
memory bandwidth: 75.2 GB/s aggregate (24 threads) · 44.3 GB/s single core (59% of bus)
strawmann startup banner, 2026-09-29T04:30:25Z; the Qdrant run agrees within 10%
strawmANN run start: environment hash 373e8531fd3d1bf6
Qdrant run start: environment hash 373e8531fd3d1bf6
All throughput rows every measurement, with its notes
workloadstrawmANNQdrantrationotes
W0d=4 floor: graph traversal with the distance taken out4,8664,0471.20x1.20x is 1.65x less work per query x 0.65x cores busy during the row (0.65 against 1.00) x 1.00x clock x 1.12x counted on-CPU share: the middle term is occupancy, not search speed d=4 floor: this measures graph traversal and heaps with the distance taken out, not request plumbing (findings 28: 76% of the server's CPU was HNSW search, 24% kernel; the RPC path is a small remainder)
W3search, fp32, single query1,2819861.30x
W4search, saturating (closed loop)4,6565,6830.82x
W4-sat50search, fixed rate at 50% of saturation (open loop)2,3282,841offered
W4-sat70search, fixed rate at 70% of saturation (open loop)3,2593,978offered
W4-sat90search, fixed rate at 90% of saturation (open loop)4,1915,115offered
W5search batched (16 distinct dataset queries per request)4,656291 requests/s2,164136 requests/s2.15xper-batch latency (16 queries/request); queries: dataset, random-sample 2.15x is 0.59x less work per query x 3.60x cores busy during the row (7.10 against 1.97) x 1.00x clock x 1.01x counted on-CPU share: the middle term is occupancy, not search speed batched 16, so qps and rps differ by 16x; the ratio uses qps
W6quantized: scalar3,6212,7261.33x
W6-ef32SQ8 recall control, ef=32 (latency only)7,6195,2681.45x1.45x is 1.76x less work per query x 0.78x cores busy during the row (1.55 against 2.00) x 1.00x clock x 1.06x counted on-CPU share: the middle term is occupancy, not search speed
W6-ef64SQ8 recall control, ef=64 (latency only)5,0583,9851.27x
W6-ef128SQ8 recall control, ef=128 (latency only)3,6152,7311.32x
W6-ef256SQ8 recall control, ef=256 (latency only)2,3971,7341.38x
W6-ef512SQ8 recall control, ef=512 (latency only)1,4231,0051.42x
W7quantized: binary + oversampling4,3524,354paritywithin the ±2.6% band this dataset's noise floor puts on W7: no measured difference, not a small one
W8quantized: PQ1,4521,1211.30x
W9exact / brute force7171paritywithin the ±3.1% band this dataset's noise floor puts on W9: no measured difference, not a small one brute force over the whole collection, no index involved
W10-ef32recall control, ef=32 (latency only)11,96710,4281.15x1.15x is 0.92x less work per query x 1.21x cores busy during the row (6.96 against 5.77) x 1.00x clock x 1.04x counted on-CPU share: the middle term is occupancy, not search speed
W10-ef64recall control, ef=64 (latency only)7,4767,1411.05x
W10-ef128recall control, ef=128 (latency only)4,6364,7910.97x
W10-ef256recall control, ef=256 (latency only)2,8523,1200.91x
W10-ef512recall control, ef=512 (latency only)1,7311,9590.88x
W12-sel1filtered search, one keyword (~1% of bench12)2,5942,503paritywithin the ±9.1% band this dataset's noise floor puts on W12-sel1: no measured difference, not a small one filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set
W12-sel10filtered search, any of 10 keywords (~10% of bench12)1,1471,280paritywithin the ±12.7% band this dataset's noise floor puts on W12-sel10: no measured difference, not a small one filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set
W12-sel1-ef32filtered recall control, one keyword, ef=32 (latency only)2,6054,5990.57x
W12-sel1-ef64filtered recall control, one keyword, ef=64 (latency only)2,5993,4860.75x
W12-sel1-ef128filtered recall control, one keyword, ef=128 (latency only)2,6122,510paritywithin the ±9.1% band this dataset's noise floor puts on W12-sel1-ef128: no measured difference, not a small one
W12-sel1-ef256filtered recall control, one keyword, ef=256 (latency only)2,5931,7421.49x
W12-sel1-ef512filtered recall control, one keyword, ef=512 (latency only)2,6061,1372.29x
W12-sel10-ef32filtered recall control, any of 10 keywords, ef=32 (latency only)2,5832,998-recall unequal: 0.9929 vs 0.8703; §7.4 compares at equal recall
W12-sel10-ef64filtered recall control, any of 10 keywords, ef=64 (latency only)1,7612,081-recall unequal: 0.9975 vs 0.9666; §7.4 compares at equal recall
W12-sel10-ef128filtered recall control, any of 10 keywords, ef=128 (latency only)1,1461,268paritywithin the ±12.7% band this dataset's noise floor puts on W12-sel10-ef128: no measured difference, not a small one
W12-sel10-ef256filtered recall control, any of 10 keywords, ef=256 (latency only)4087840.52x
W12-sel10-ef512filtered recall control, any of 10 keywords, ef=512 (latency only)4144860.85x
W13scroll / pagination50,84546,4501.09xstrawmANN: -n 200000 (4x the table's 50000) Qdrant: -n 100000 (2x the table's 50000) strawmANN: 0.13 ms of the client's 0.14 ms p50 is not server time (13x), so 92% of this row is the load generator and the socket Qdrant: 0.13 ms of the client's 0.16 ms p50 is not server time (6x), so 82% of this row is the load generator and the socket 1.09x is 3.29x less work per query x 0.33x cores busy during the row (1.39 against 4.24) x 1.00x clock x 1.02x counted on-CPU share: the middle term is occupancy, not search speed
W11-steadymixed read/write below the rebuild threshold: search bench2 while 5,255 synthetic points append2,3883,857-strawmANN: search covered 79% of the append; append 198 points/s Qdrant: search covered 49% of the append; append 198 points/s search-during-write; no recall join
W11mixed read/write: search bench2 while 21,020 synthetic points append (runs last)6233,453-strawmANN: write overlap 88%; append 299 points/s Qdrant: search covered 21% of the append; append 299 points/s search-during-write; no recall join; the writer covered under 90% of strawmANN's search, so the last 12% of that row measured the rebuild the append provoked rather than a concurrent write

W10: throughput against ef

ef is the same unit here: Qdrant held 2 segments for this run, one populated and 1 empty (read back per segment from the engine), and strawmANN held one graph, so equal ef is equal traversal width and the curves may be read at equal x. That is a property of this run, not of the engines.
All latency rows closed loop included, p50 to max

A row flagged “… not server time” is one where the client saw far more than the server reported: below the rate at which a queue can form, what is left is the load generator's own scheduling.

workloadstrawmANN p50strawmANN p95strawmANN p99strawmANN p99.9strawmANN maxQdrant p50Qdrant p95Qdrant p99Qdrant p99.9Qdrant max
W0d=4 floor: graph traversal with the distance taken out closed loop 201 µs230 µs249 µs315 µs1.41 ms243 µs275 µs303 µs374 µs1.41 ms
W3search, fp32, single query closed loop 759 µs1.07 ms1.40 ms1.73 ms4.02 ms1.00 ms1.32 ms1.43 ms1.54 ms3.50 ms
W4search, saturating (closed loop) closed loop 13.69 ms14.80 ms15.28 ms34.24 ms44.95 ms10.97 ms16.15 ms22.55 ms35.02 ms54.93 ms
W4-sat50search, fixed rate at 50% of saturation (open loop) open loop 889 µs1.34 ms1.57 ms1.80 ms5.50 ms1.12 ms1.54 ms2.09 ms2.49 ms9.21 ms
W4-sat70search, fixed rate at 70% of saturation (open loop) open loop 1.01 ms1.80 ms2.12 ms2.47 ms6.13 ms1.30 ms2.27 ms2.65 ms3.56 ms10.24 ms
W4-sat90search, fixed rate at 90% of saturation (open loop) open loop 1.15 ms1.99 ms2.57 ms3.51 ms7.14 ms1.91 ms4.45 ms8.26 ms13.36 ms22.75 ms
W5search batched (16 distinct dataset queries per request) closed loop 6.86 ms7.77 ms8.12 ms8.55 ms14.04 ms14.78 ms16.05 ms16.64 ms17.30 ms18.72 ms
W6quantized: scalar closed loop 543 µs672 µs842 µs1.12 ms1.67 ms728 µs876 µs938 µs1.03 ms3.50 ms
W6-ef32SQ8 recall control, ef=32 (latency only) closed loop 240 µs387 µs435 µs509 µs1.17 ms375 µs432 µs464 µs563 µs3.71 ms
W6-ef64SQ8 recall control, ef=64 (latency only) closed loop 377 µs580 µs700 µs793 µs1.39 ms498 µs580 µs618 µs716 µs1.92 ms
W6-ef128SQ8 recall control, ef=128 (latency only) closed loop 543 µs674 µs860 µs1.14 ms1.71 ms726 µs875 µs936 µs1.02 ms3.77 ms
W6-ef256SQ8 recall control, ef=256 (latency only) closed loop 826 µs1.02 ms1.10 ms1.18 ms2.28 ms1.14 ms1.42 ms1.52 ms1.63 ms3.14 ms
W6-ef512SQ8 recall control, ef=512 (latency only) closed loop 1.39 ms1.74 ms1.86 ms2.01 ms3.65 ms1.98 ms2.46 ms2.63 ms2.80 ms4.63 ms
W7quantized: binary + oversampling closed loop 433 µs614 µs677 µs841 µs1.25 ms455 µs503 µs536 µs651 µs1.81 ms
W8quantized: PQ closed loop 1.36 ms1.74 ms1.93 ms2.28 ms3.05 ms1.76 ms2.10 ms2.26 ms2.57 ms3.91 ms
W9exact / brute force closed loop 109.30 ms147.44 ms164.37 ms180.92 ms181.57 ms113.19 ms122.09 ms128.73 ms140.29 ms146.54 ms
W10-ef32recall control, ef=32 (latency only) closed loop 658 µs866 µs983 µs1.22 ms3.51 ms714 µs1.17 ms1.39 ms1.80 ms3.92 ms
W10-ef64recall control, ef=64 (latency only) closed loop 1.05 ms1.45 ms1.64 ms1.87 ms6.02 ms1.03 ms1.75 ms2.06 ms2.63 ms6.11 ms
W10-ef128recall control, ef=128 (latency only) closed loop 1.69 ms2.43 ms2.76 ms3.19 ms7.54 ms1.56 ms2.58 ms3.01 ms3.83 ms7.27 ms
W10-ef256recall control, ef=256 (latency only) closed loop 2.75 ms4.06 ms4.63 ms5.37 ms8.58 ms2.46 ms3.87 ms4.57 ms5.64 ms9.31 ms
W10-ef512recall control, ef=512 (latency only) closed loop 4.52 ms6.79 ms7.75 ms8.93 ms11.22 ms3.99 ms6.02 ms7.07 ms8.50 ms11.40 ms
W12-sel1filtered search, one keyword (~1% of bench12) closed loop 767 µs819 µs864 µs930 µs2.24 ms797 µs903 µs949 µs1.03 ms2.83 ms
W12-sel10filtered search, any of 10 keywords (~10% of bench12) closed loop 1.74 ms2.29 ms2.48 ms2.68 ms4.41 ms1.57 ms1.86 ms1.96 ms2.08 ms4.21 ms
W12-sel1-ef32filtered recall control, one keyword, ef=32 (latency only) closed loop 764 µs815 µs861 µs970 µs2.29 ms433 µs495 µs529 µs642 µs1.75 ms
W12-sel1-ef64filtered recall control, one keyword, ef=64 (latency only) closed loop 765 µs817 µs859 µs934 µs2.31 ms574 µs654 µs692 µs795 µs2.13 ms
W12-sel1-ef128filtered recall control, one keyword, ef=128 (latency only) closed loop 762 µs810 µs848 µs939 µs2.27 ms795 µs901 µs944 µs1.02 ms2.67 ms
W12-sel1-ef256filtered recall control, one keyword, ef=256 (latency only) closed loop 766 µs820 µs863 µs936 µs2.30 ms1.14 ms1.26 ms1.32 ms1.40 ms3.35 ms
W12-sel1-ef512filtered recall control, one keyword, ef=512 (latency only) closed loop 763 µs815 µs857 µs936 µs2.28 ms1.76 ms1.89 ms1.95 ms2.07 ms4.18 ms
W12-sel10-ef32filtered recall control, any of 10 keywords, ef=32 (latency only) closed loop 767 µs972 µs1.05 ms1.24 ms2.18 ms661 µs791 µs862 µs965 µs2.34 ms
W12-sel10-ef64filtered recall control, any of 10 keywords, ef=64 (latency only) closed loop 1.13 ms1.47 ms1.59 ms1.75 ms2.96 ms963 µs1.13 ms1.20 ms1.31 ms3.70 ms
W12-sel10-ef128filtered recall control, any of 10 keywords, ef=128 (latency only) closed loop 1.75 ms2.29 ms2.49 ms2.70 ms4.52 ms1.59 ms1.87 ms1.97 ms2.09 ms4.85 ms
W12-sel10-ef256filtered recall control, any of 10 keywords, ef=256 (latency only) closed loop 4.88 ms5.36 ms5.61 ms5.90 ms10.05 ms2.59 ms3.03 ms3.18 ms3.34 ms5.95 ms
W12-sel10-ef512filtered recall control, any of 10 keywords, ef=512 (latency only) closed loop 4.81 ms5.18 ms5.34 ms5.67 ms9.19 ms4.17 ms4.91 ms5.16 ms5.41 ms8.57 ms
W13scroll / pagination closed loop 138 µs205 µs232 µs276 µs1.08 ms158 µs223 µs304 µs466 µs1.51 ms
W11-steadymixed read/write below the rebuild threshold: search bench2 while 5,255 synthetic points append closed loop 3.10 ms5.72 ms6.75 ms7.80 ms9.31 ms1.83 ms3.54 ms5.33 ms16.33 ms32.41 ms
W11mixed read/write: search bench2 while 21,020 synthetic points append (runs last) closed loop 10.83 ms23.04 ms27.03 ms31.01 ms35.19 ms2.07 ms3.95 ms6.82 ms17.67 ms31.98 ms
Storage and I/O block layer and syscalls, per row

The syscall rows count every descriptor, sockets included, so on a search row they measure the network rather than the disk.

strawmANNQdrant
storage on disk6.0 GiBbefore W11-steady, W114.6 GiBbefore W11-steady, W11
peak RSS5.6 GiB23.8 GiB8.6 GiB before the writers
of which anonymous609.2 MiB550.4 MiB
of which file-backed5.0 GiB7.1 GiB
disk read bytes20.0 KiB282.4 MiB
disk write bytes5.0 GiB53.9 GiB
disk read ops158,733
disk write ops82,932903,863
syscall reads (all fds)10,275,00910,196
syscall writes (all fds)4,599,95211,471,461

measured via strawmANN: proc+cgroup / Qdrant: proc+cgroup. proc supplies syscall counts and block-layer bytes, cgroup supplies block-layer operations and bytes, so a row one interface does not carry reads n/a via …. unknown means it was not measured, and 0 means it was: an engine started with no --data-dir has no store, which is the row this section exists for. storage on disk is the level before the rows with a concurrent writer: during those an engine that rewrites segments is caught mid-rewrite, and the same row has read 3.47 and 10.11 GiB on two runs of one binary. peak RSS does include them, being a peak.

Per workload

A search row doing block-layer reads is an engine going to disk to answer a query.

workloadstrawmANN readstrawmANN writtenstrawmANN read opsstrawmANN write opsQdrant readQdrant writtenQdrant read opsQdrant write ops
W0-upload06.5 MiB0010.8 MiB36.9 MiB4531,320
W00001050000
W10821.1 MiB00444.0 KiB2.2 GiB4225,045
W20821.1 MiB026,50212.0 MiB3.7 GiB45663,640
W300030006
W400000000
W4-sat5000000000
W4-sat7000000000
W4-sat9000000000
W500000000
W6-upload0821.1 MiB013,24614.2 MiB4.6 GiB50178,766
W600030006
W6-ef3200000000
W6-ef6400000000
W6-ef12800000000
W6-ef25600000000
W6-ef51200000000
W7-upload0821.1 MiB013,24315.5 MiB4.5 GiB58876,196
W700030006
W8-upload8.0 KiB821.1 MiB613,24613.5 MiB3.6 GiB40861,481
W800030000
W900000000
W10-ef3200000000
W10-ef6400000000
W10-ef12800000000
W10-ef25600000000
W10-ef51200000000
W12-upload12.0 KiB821.1 MiB913,24324.6 MiB3.7 GiB80764,840
W12-sel100030000
W12-sel1000000000
W12-sel1-ef3200000000
W12-sel1-ef6400000000
W12-sel1-ef12800000000
W12-sel1-ef25600000000
W12-sel1-ef51200000000
W12-sel10-ef3200000000
W12-sel10-ef6400000000
W12-sel10-ef12800000000
W12-sel10-ef25600000000
W12-sel10-ef51200000000
W1300000000
W11-steady041.2 MiB0038.3 MiB5.9 GiB1,089100,674
W110164.5 MiB03,332152.9 MiB25.7 GiB4,389431,883
Scheduler and memory CPU use, run-queue wait, migrations

Whether the engine was running while it ran. of wall is CPU over elapsed, waiting is runnable-but-not-scheduled, migrations checks the pinning claim, and switches gives voluntary over involuntary — the scheduler taking the core away against the engine choosing to sleep, which per unit of work is the cheapest signal of lock contention there is. A dash is not a zero: an index build's threads exit before the row does, and threads shows what fraction of the row the survivors account for.

Time spent waiting for a core

Lower is better; the bar between a pair is the gap.

Peak memory per row

Peak RSS while the row ran, from the engine's own process. Unlike the disk counters this is not refused across the two engines: residency decides where bytes live, and this is what the process held either way.

workloadstrawmANNQdrant
cpuof wallwaitingswitches vol/involmigrationsfaults min/majthreadscpuof wallwaitingswitches vol/involmigrationsfaults min/majthreads
W0-upload9.0 s146%4.2 ms4,553 / 36013,130 / 09 2%12.4 s187%1,180.2 ms17,205 / 2,0321,91049,093 / 127,99242 14%
W07.1 s68%4.4 ms164,008 / 1100 / 0912.6 s102%3.7 ms560,500 / 16149282 / 037 0%
W12.0 s64%1.0 ms97,353 / 30225,897 / 0918.2 s390%7,564.6 ms19,232 / 5,7684,367160,649 / 251,72154 72%
W2167.8 s639%99.0 ms98,924 / 8300229,765 / 09 1%144.6 s529%7,150.1 ms25,230 / 6,4535,343261,652 / 321,68143 7%
W335.6 s91%3.3 ms191,148 / 5405 / 0951.7 s102%4.9 ms564,991 / 1151741,802 / 039 0%
W477.3 s718%7.7 ms82,992 / 104084 / 0969.6 s788%67,279.5 ms42,137 / 37,5937,0235,372 / 076
W4-sat50167.8 s195%24.3 ms489,583 / 116012 / 09224.8 s319%22,057.5 ms1,696,973 / 13,810497,1753,260 / 068 0%
W4-sat70196.9 s320%26.2 ms446,669 / 135020 / 09244.4 s485%93,503.0 ms1,383,575 / 77,473583,3292,241 / 068
W4-sat90222.4 s465%29.4 ms401,908 / 171048 / 09259.1 s660%161,417.8 ms713,168 / 147,070267,5142,445 / 044 0%
W576.3 s710%7.2 ms6,624 / 14500 / 0945.7 s198%9.6 ms38,060 / 3381,2801,284 / 042 0%
W6-upload227.9 s258%107.3 ms29,404 / 1,3773716,129 / 09 1%85.1 s354%3,386.9 ms30,152 / 3,2133,619432,188 / 355,58742 0%
W624.1 s174%6.3 ms184,979 / 4100 / 0937.2 s203%10.9 ms530,264 / 11114,687985 / 044 0%
W6-ef3210.2 s155%2.2 ms166,239 / 600 / 0919.2 s202%8.5 ms493,652 / 389,455122 / 044
W6-ef6416.3 s164%3.8 ms180,002 / 1100 / 0925.5 s203%9.1 ms513,176 / 4211,474129 / 044
W6-ef12824.2 s175%2.9 ms184,479 / 1800 / 0937.3 s203%10.8 ms530,140 / 11214,753476 / 044
W6-ef25638.2 s183%4.2 ms183,644 / 2500 / 0958.5 s202%12.9 ms541,209 / 20918,679640 / 044
W6-ef51266.7 s190%6.3 ms186,701 / 10304 / 09100.7 s202%71.4 ms548,977 / 43820,7761,376 / 044
W7-upload167.5 s592%502.4 ms26,300 / 9775237,989 / 09 1%37.8 s235%2,105.4 ms30,544 / 1,8083,933328,498 / 413,12042 0%
W719.5 s169%5.8 ms183,976 / 3300 / 0923.1 s201%12.1 ms486,770 / 358,960506 / 046 0%
W8-upload1,362.4 s771%433.9 ms31,987 / 6,4381461,956 / 09 0%955.9 s636%2,287.5 ms329,850 / 14,65015,664252,009 / 287,84342 0%
W865.3 s189%4.9 ms189,322 / 8100 / 0989.9 s202%136.5 ms520,095 / 46920,3255,327 / 046
W9197.4 s702%28.8 ms4,607 / 40500 / 09221.5 s788%11,013.9 ms20,079 / 11,8558,603787 / 048
W10-ef3228.9 s686%4.6 ms96,554 / 4200 / 0927.9 s576%10,904.3 ms241,833 / 21,48572,1621,315 / 048
W10-ef6447.1 s701%5.6 ms100,740 / 5900 / 0943.5 s617%28,791.1 ms290,429 / 38,056129,916973 / 070
W10-ef12876.5 s707%9.1 ms104,056 / 11204 / 0966.2 s632%41,856.1 ms313,068 / 41,854150,106577 / 063 0%
W10-ef256124.4 s708%13.4 ms108,762 / 21304 / 09105.4 s656%58,014.1 ms333,899 / 55,939178,287724 / 065
W10-ef512204.1 s706%22.7 ms112,878 / 37900 / 09174.0 s681%80,763.0 ms363,425 / 80,921197,4521,495 / 065
W12-upload166.8 s634%94.9 ms21,802 / 8121230,557 / 09 1%219.2 s436%7,688.9 ms37,271 / 8,8464,185214,691 / 352,82650 0%
W12-sel134.8 s180%3.0 ms160,548 / 6700 / 0940.6 s203%10.9 ms529,296 / 11313,621673 / 055
W12-sel1083.3 s191%4.9 ms186,419 / 12100 / 0979.1 s202%48.9 ms545,876 / 35720,685712 / 055
W12-sel1-ef3234.7 s180%3.2 ms159,742 / 1900 / 0922.1 s202%8.8 ms503,380 / 6210,523204 / 055
W12-sel1-ef6434.7 s180%5.1 ms159,303 / 3700 / 0929.2 s203%12.2 ms516,391 / 9412,016289 / 055
W12-sel1-ef12834.6 s180%2.1 ms161,430 / 4100 / 0940.4 s203%13.2 ms528,221 / 15014,175292 / 055
W12-sel1-ef25634.7 s179%1.9 ms159,960 / 3400 / 0958.1 s202%20.9 ms529,766 / 22914,920593 / 055
W12-sel1-ef51234.6 s180%5.1 ms161,106 / 2200 / 0988.2 s200%117.3 ms511,258 / 49013,954942 / 055
W12-sel10-ef3235.2 s182%9.3 ms183,179 / 3500 / 0933.8 s202%10.3 ms524,793 / 13414,208538 / 055
W12-sel10-ef6453.1 s187%3.9 ms183,905 / 5500 / 0948.7 s202%16.8 ms534,062 / 19216,801266 / 055
W12-sel10-ef12883.5 s191%5.6 ms186,655 / 15700 / 0979.8 s202%62.9 ms544,704 / 35720,657940 / 055
W12-sel10-ef256231.5 s189%31.8 ms115,826 / 94400 / 09128.6 s202%76.6 ms554,834 / 46519,9931,458 / 055
W12-sel10-ef512231.3 s192%35.0 ms113,692 / 88800 / 09204.8 s199%87.0 ms571,091 / 1,56622,1742,574 / 055
W135.6 s137%2.4 ms243,523 / 1100 / 099.2 s414%932.5 ms381,150 / 4,571112,908425 / 068
W11-steady147.1 s701%20.7 ms105,734 / 297012,338 / 0996.9 s745%77,335.9 ms349,636 / 75,751171,847403,201 / 130,78646 22%
W11606.7 s756%309,613.4 ms110,900 / 57,2717158,385 / 09 87%156.9 s1,081%100,964.5 ms407,895 / 96,883182,0171,670,641 / 591,14554 35%
Waiting and stalls where the time went when the engine was not running

waiting is runnable and not scheduled, blocked on disk is not runnable at all, and pressure is PSI for the engine's cgroup (some = at least one task stalled, full = every runnable task).

Time stalled on I/O

Time every task in the engine's cgroup was stalled on I/O. Lower is better. 34 rows never stalled on I/O and are not drawn.

workloadstrawmANNQdrant
waitingblocked on diskcpu some/fullio some/fullmemory some/fullwaitingblocked on diskcpu some/fullio some/fullmemory some/full
W0-upload4.2 ms0 ms2 / 2 ms0 / 0 ms0 / 0 ms1,180.2 ms0 ms118 / 3 ms13 / 13 ms0 / 0 ms
W04.4 ms0 ms49 / 49 ms0 / 0 ms0 / 0 ms3.7 ms0 ms70 / 70 ms0 / 0 ms0 / 0 ms
W11.0 ms0 ms8 / 8 ms2 / 2 ms0 / 0 ms7,564.6 ms0 ms734 / 5 ms57 / 29 ms0 / 0 ms
W299.0 ms0 ms20 / 11 ms1 / 1 ms0 / 0 ms7,150.1 ms0 ms664 / 10 ms149 / 105 ms0 / 0 ms
W33.3 ms0 ms42 / 42 ms0 / 0 ms0 / 0 ms4.9 ms0 ms97 / 97 ms0 / 0 ms0 / 0 ms
W47.7 ms0 ms4 / 4 ms0 / 0 ms0 / 0 ms67,279.5 ms0 ms3,490 / 2 ms0 / 0 ms0 / 0 ms
W4-sat5024.3 ms0 ms160 / 160 ms0 / 0 ms0 / 0 ms22,057.5 ms0 ms2,623 / 219 ms0 / 0 ms0 / 0 ms
W4-sat7026.2 ms0 ms148 / 148 ms0 / 0 ms0 / 0 ms93,503.0 ms0 ms9,318 / 146 ms0 / 0 ms0 / 0 ms
W4-sat9029.4 ms0 ms117 / 117 ms0 / 0 ms0 / 0 ms161,417.8 ms0 ms12,693 / 71 ms0 / 0 ms0 / 0 ms
W57.2 ms0 ms3 / 3 ms0 / 0 ms0 / 0 ms9.6 ms0 ms15 / 15 ms0 / 0 ms0 / 0 ms
W6-upload107.3 ms0 ms33 / 20 ms0 / 0 ms0 / 0 ms3,386.9 ms0 ms331 / 13 ms359 / 349 ms0 / 0 ms
W66.3 ms0 ms40 / 40 ms0 / 0 ms0 / 0 ms10.9 ms0 ms63 / 62 ms0 / 0 ms0 / 0 ms
W6-ef322.2 ms0 ms24 / 24 ms0 / 0 ms0 / 0 ms8.5 ms0 ms52 / 52 ms0 / 0 ms0 / 0 ms
W6-ef643.8 ms0 ms39 / 39 ms0 / 0 ms0 / 0 ms9.1 ms0 ms57 / 56 ms0 / 0 ms0 / 0 ms
W6-ef1282.9 ms0 ms40 / 40 ms0 / 0 ms0 / 0 ms10.8 ms0 ms63 / 62 ms0 / 0 ms0 / 0 ms
W6-ef2564.2 ms0 ms38 / 38 ms0 / 0 ms0 / 0 ms12.9 ms0 ms75 / 74 ms0 / 0 ms0 / 0 ms
W6-ef5126.3 ms0 ms39 / 39 ms0 / 0 ms0 / 0 ms71.4 ms0 ms136 / 128 ms0 / 0 ms0 / 0 ms
W7-upload502.4 ms0 ms69 / 6 ms8 / 8 ms0 / 0 ms2,105.4 ms0 ms271 / 11 ms198 / 173 ms0 / 0 ms
W75.8 ms0 ms34 / 34 ms0 / 0 ms0 / 0 ms12.1 ms0 ms55 / 54 ms0 / 0 ms0 / 0 ms
W8-upload433.9 ms0 ms57 / 11 ms3 / 3 ms0 / 0 ms2,287.5 ms0 ms917 / 120 ms256 / 153 ms0 / 0 ms
W84.9 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms136.5 ms0 ms122 / 106 ms0 / 0 ms0 / 0 ms
W928.8 ms0 ms4 / 4 ms0 / 0 ms0 / 0 ms11,013.9 ms0 ms1,259 / 21 ms0 / 0 ms0 / 0 ms
W10-ef324.6 ms0 ms12 / 12 ms0 / 0 ms0 / 0 ms10,904.3 ms0 ms1,266 / 21 ms0 / 0 ms0 / 0 ms
W10-ef645.6 ms0 ms11 / 11 ms0 / 0 ms0 / 0 ms28,791.1 ms0 ms2,667 / 24 ms0 / 0 ms0 / 0 ms
W10-ef1289.1 ms0 ms10 / 10 ms0 / 0 ms0 / 0 ms41,856.1 ms0 ms3,867 / 26 ms0 / 0 ms0 / 0 ms
W10-ef25613.4 ms0 ms9 / 9 ms0 / 0 ms0 / 0 ms58,014.1 ms0 ms5,612 / 35 ms0 / 0 ms0 / 0 ms
W10-ef51222.7 ms0 ms9 / 9 ms0 / 0 ms0 / 0 ms80,763.0 ms0 ms8,206 / 39 ms0 / 0 ms0 / 0 ms
W12-upload94.9 ms0 ms15 / 6 ms10 / 10 ms0 / 0 ms7,688.9 ms0 ms758 / 18 ms148 / 101 ms0 / 0 ms
W12-sel13.0 ms0 ms38 / 38 ms0 / 0 ms0 / 0 ms10.9 ms0 ms64 / 63 ms0 / 0 ms0 / 0 ms
W12-sel104.9 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms48.9 ms0 ms112 / 107 ms0 / 0 ms0 / 0 ms
W12-sel1-ef323.2 ms0 ms38 / 38 ms0 / 0 ms0 / 0 ms8.8 ms0 ms54 / 54 ms0 / 0 ms0 / 0 ms
W12-sel1-ef645.1 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms12.2 ms0 ms59 / 58 ms0 / 0 ms0 / 0 ms
W12-sel1-ef1282.1 ms0 ms38 / 38 ms0 / 0 ms0 / 0 ms13.2 ms0 ms64 / 63 ms0 / 0 ms0 / 0 ms
W12-sel1-ef2561.9 ms0 ms38 / 38 ms0 / 0 ms0 / 0 ms20.9 ms0 ms72 / 71 ms0 / 0 ms0 / 0 ms
W12-sel1-ef5125.1 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms117.3 ms0 ms108 / 94 ms0 / 0 ms0 / 0 ms
W12-sel10-ef329.3 ms0 ms36 / 36 ms0 / 0 ms0 / 0 ms10.3 ms0 ms61 / 61 ms0 / 0 ms0 / 0 ms
W12-sel10-ef643.9 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms16.8 ms0 ms68 / 67 ms0 / 0 ms0 / 0 ms
W12-sel10-ef1285.6 ms0 ms37 / 37 ms0 / 0 ms0 / 0 ms62.9 ms0 ms114 / 108 ms0 / 0 ms0 / 0 ms
W12-sel10-ef25631.8 ms0 ms42 / 42 ms0 / 0 ms0 / 0 ms76.6 ms0 ms144 / 139 ms0 / 0 ms0 / 0 ms
W12-sel10-ef51235.0 ms0 ms42 / 42 ms0 / 0 ms0 / 0 ms87.0 ms0 ms162 / 156 ms0 / 0 ms0 / 0 ms
W132.4 ms0 ms28 / 28 ms0 / 0 ms0 / 0 ms932.5 ms0 ms130 / 26 ms0 / 0 ms0 / 0 ms
W11-steady20.7 ms0 ms15 / 15 ms0 / 0 ms0 / 0 ms77,335.9 ms0 ms5,714 / 49 ms344 / 323 ms0 / 0 ms
W11309,613.4 ms0 ms38,673 / 19 ms10 / 6 ms0 / 0 ms100,964.5 ms0 ms6,956 / 71 ms2,067 / 2,014 ms0 / 0 ms

A dash under blocked on disk is not a zero: most hosts ship with kernel.task_delayacct=0, which reports the field as a permanent zero and would manufacture the strongest claim here — that the memory-resident engine never waits for a device — out of a sysctl. bench/setup.py apply turns it on.

Hardware counters cycles, DRAM loads, TLB walks, IPC per query

What the core did per query, from perf stat attached to the engine for the same window as the /proc counters. §5's cost model is a set of claims about instructions per cycle, DRAM traffic and TLB reach; these are those quantities, measured on the row the headline quotes rather than on a microbenchmark.

What one query cost, in cycles

Lower is better; the bar between a pair is the gap. Load rows are absent: they have no queries to divide by, and an absolute count under a per-query axis would be a different quantity wearing this one's label.

Demand loads from DRAM per query

Demand loads served from DRAM, times the cache line this host reports. Lower is better. Demand only: the hardware prefetcher's fills are a separate counter and are not in this number, so a row that streams — an exact scan above all — moved far more than this says. On the graph rows, where there is little for a prefetcher to predict, it is most of the traffic and is the outstanding-miss quantity §5.2 argues the engine is limited by.

TLB walks per query

Data TLB misses that reached a page walk, per query. Lower is better. §5.5 argues a memory-resident index needs hugepage care; this is what not taking it costs, and `--no-huge-pages` is the A/B arm that prices it.

Instructions per cycle

How well the core was fed while it ran. Not a score: a low IPC on a memory-bound row is what §5.2 predicts, and a high one on a row that does more work per query is not a win. Read it beside the two charts above.

Branch mispredictions per 1k instructions

A graph traversal is a chain of data-dependent branches and none of them is predictable from the last query, so this is the column that separates a slow kernel from an unpredictable walk. Per thousand instructions rather than per branch: `branches` costs a PMU counter that `ref-cycles` needs more, and MPKI is the comparable form anyway.

workloadstrawmANNQdrant
IPCGHzbranch MPKIcycles/querydemand DRAM/queryTLB walks/queryIPCGHzbranch MPKIcycles/querydemand DRAM/queryTLB walks/query
W0-upload1.932.001.004x nominal5.40---1.842.011.005x nominal4.53---
W01.342.011.005x nominal6.05229,5514.4 KiB633.91.412.011.004x nominal3.12378,7577.1 KiB699.8
W11.822.011.005x nominal0.61---1.692.001.003x nominal0.89---
W20.422.011.005x nominal2.37---0.732.011.005x nominal1.72---
W30.692.011.005x nominal4.011,343,8063,937.2 KiB2,447.60.822.011.006x nominal2.221,926,1852,248.8 KiB4,407.3
W40.302.011.006x nominal3.263,086,3002,809.6 KiB2,346.10.532.011.006x nominal2.322,743,4632,112.1 KiB3,603.8
W4-sat500.572.011.005x nominal4.051,625,4843,554.2 KiB2,363.90.742.011.006x nominal2.522,107,2552,263.7 KiB4,316.8
W4-sat700.472.011.006x nominal3.921,951,2563,262.1 KiB2,338.40.682.011.005x nominal2.572,302,9372,224.5 KiB4,170.7
W4-sat900.422.011.005x nominal3.672,199,3432,925.4 KiB2,311.10.602.011.006x nominal2.542,504,9202,119.5 KiB3,894.9
W50.282.011.005x nominal3.443,053,3572,801.5 KiB2,285.40.722.011.006x nominal2.551,813,7182,334.2 KiB3,468.9
W6-upload0.592.011.005x nominal12.01---1.902.011.005x nominal1.59---
W60.922.011.005x nominal4.81897,6682,476.9 KiB2,729.61.512.011.006x nominal1.901,368,830681.3 KiB2,990.0
W6-ef321.012.011.006x nominal2.53373,764748.2 KiB1,162.11.682.011.006x nominal1.41656,743231.2 KiB1,449.4
W6-ef640.902.011.005x nominal4.16589,8331,411.1 KiB1,769.81.612.011.006x nominal1.63903,927387.5 KiB2,013.4
W6-ef1280.922.011.005x nominal4.79899,6952,475.3 KiB2,729.11.512.011.006x nominal1.901,366,164680.2 KiB2,990.1
W6-ef2560.942.011.005x nominal5.291,463,0014,321.8 KiB4,436.51.422.011.006x nominal2.162,206,7421,230.5 KiB4,592.3
W6-ef5120.932.011.006x nominal5.632,603,8407,520.3 KiB7,889.91.322.011.006x nominal2.593,825,5542,295.7 KiB7,272.7
W7-upload0.422.011.005x nominal2.94---2.032.011.005x nominal4.55---
W70.852.011.006x nominal9.15724,084726.7 KiB1,811.31.462.011.005x nominal3.60811,466306.7 KiB2,155.0
W8-upload6.262.001.004x nominal0.48---5.572.011.004x nominal0.16---
W83.942.011.006x nominal0.502,546,739620.5 KiB2,848.23.202.011.006x nominal0.553,437,700429.7 KiB3,954.5
W90.242.001.004x nominal0.23197,681,596236,938.1 KiB65,509.70.362.011.006x nominal0.14221,203,34682,785.0 KiB212,349.5
W10-ef320.362.011.006x nominal2.231,139,494981.0 KiB815.40.782.011.006x nominal2.061,047,142846.8 KiB1,682.6
W10-ef640.322.011.006x nominal2.811,869,6691,680.8 KiB1,347.00.682.011.006x nominal2.351,627,9431,340.7 KiB2,553.9
W10-ef1280.302.011.006x nominal3.303,046,4922,840.4 KiB2,342.90.632.011.006x nominal2.572,508,0092,155.5 KiB3,954.2
W10-ef2560.302.011.006x nominal3.794,961,6134,791.7 KiB4,257.10.582.011.006x nominal2.824,045,3423,435.1 KiB6,208.7
W10-ef5120.312.011.006x nominal4.118,166,3908,023.1 KiB8,115.60.552.011.006x nominal2.966,760,9675,728.7 KiB9,967.1
W12-upload0.422.011.006x nominal2.40---0.862.001.004x nominal1.82---
W12-sel10.502.011.006x nominal2.671,328,6174,467.9 KiB2,131.31.002.011.006x nominal3.441,497,8891,218.6 KiB2,347.9
W12-sel101.102.011.006x nominal4.133,277,4454,930.4 KiB6,298.60.942.011.006x nominal2.612,993,7182,705.6 KiB3,681.8
W12-sel1-ef320.502.011.006x nominal2.651,326,9744,470.3 KiB2,130.71.112.011.006x nominal2.33767,892563.2 KiB1,488.3
W12-sel1-ef640.502.011.006x nominal2.691,329,0054,470.3 KiB2,133.01.042.011.006x nominal2.871,047,413826.2 KiB1,878.0
W12-sel1-ef1280.502.011.006x nominal2.661,322,8264,466.8 KiB2,130.81.002.011.006x nominal3.451,494,5911,214.2 KiB2,349.7
W12-sel1-ef2560.502.011.006x nominal2.661,329,2944,474.8 KiB2,133.01.002.011.005x nominal4.092,193,4531,726.2 KiB2,898.7
W12-sel1-ef5120.502.011.006x nominal2.651,325,4704,466.1 KiB2,129.50.992.011.006x nominal4.933,384,9162,425.0 KiB3,411.3
W12-sel10-ef321.042.011.006x nominal3.541,343,1062,070.9 KiB2,512.71.052.011.006x nominal1.971,232,8431,007.6 KiB2,000.8
W12-sel10-ef641.062.011.005x nominal3.892,060,5323,204.8 KiB3,854.61.012.011.006x nominal2.131,822,1591,610.2 KiB2,669.4
W12-sel10-ef1281.102.011.005x nominal4.143,281,9304,932.2 KiB6,303.50.932.011.006x nominal2.643,018,8622,734.5 KiB3,682.8
W12-sel10-ef2560.552.011.006x nominal0.739,224,48526,914.0 KiB15,346.90.922.011.005x nominal2.914,936,6014,436.2 KiB5,116.2
W12-sel10-ef5120.552.011.006x nominal0.799,207,19027,014.5 KiB15,339.90.942.011.005x nominal3.027,972,8656,902.6 KiB7,207.1
W131.362.011.007x nominal1.2042,4571.0 KiB12.91.452.011.006x nominal1.46139,5441.4 KiB37.1
W11-steady0.332.011.005x nominal1.755,868,5265,842.3 KiB3,195.90.792.001.004x nominal1.863,660,3942,818.8 KiB4,559.9
W110.262.011.005x nominal1.0924,190,69128,713.5 KiB9,694.31.062.011.005x nominal1.286,049,8584,563.7 KiB6,796.6

Nothing here is scaled. Where perf had to multiplex the group, the values it prints are extrapolations from the fraction of the row each counter was on, and they are withheld rather than shown — the same rule the scheduler table applies to a thread that exited. An event this host does not implement is likewise blank, never zero.

The two instruments disagree on these rows. Qdrant: perf and /proc disagree on minor faults by up to 48% on W0-upload, W2, W6-upload, W7-upload, W8-upload, W12-upload, W11-steady, W11; Qdrant: perf and /proc disagree on context switches by up to 75% on W0-upload, W1, W6-upload, W7-upload, W8-upload, W12-upload. perf keeps an exiting thread's counts and the /proc sums do not, so a gap here is usually the same lost-thread effect the scheduler table reports as coverage.

Collection config what each engine says it built

Read back from the engine rather than taken from the request. --segments 1 sets Qdrant's default_segment_number, which its optimizer treats as a target, and the segment count is the largest confound in the ef comparison. strawmANN held 1 segment; Qdrant held 2 segments.

Where the engines disagree a cell reads strawmANN / Qdrant, and is marked.

collectionsegmentspopulated segmentsrequested segmentspointsindexed vectorsvector sizehnsw mhnsw ef_constructquantizationshardsvector residency
bench01 / 2- / 1- / 1105,100105,100416100-1cached
bench11 / 6- / 6- / 1105,100 / 103,4000 / 12,7002,04816100-1cached
bench21 / 2- / 1- / 1105,100105,1002,04816100-1cached
bench61 / 2- / 1- / 1105,100105,1002,04816100scalar1cached
bench71 / 2- / 1- / 1105,100105,1002,04816100binary1cached
bench81 / 2- / 1- / 1105,100105,1002,04816100product1cached
bench121 / 2- / 1- / 1105,100105,1002,04816100-1cached

After the mutating rows bench2 held strawmANN 131,375 (+26,275), 127,843 of them indexed; Qdrant 131,375 (+26,275), 131,375 of them indexed. The table above is the state the throughput, latency and recall rows searched; this is what W11 left, and is that row's subject rather than theirs.

bench1 was read back straight after its row and dropped, with no settle between, so an engine still building it reports a snapshot taken mid-ingest. Its cells are not marked as a disagreement.

Per workload every metric of every row, engine beside engine

Metric down, engine across, so a comparison is two adjacent cells. Metrics a row did not measure are dropped rather than shown empty.

W0-upload transport floor: load

upload · collection=bench0 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=105,100
metricstrawmANNQdrant
wall clock6.2 s6.7 s1.08x
of which upload0.1 s0.6 s4.53x
of which index wait4.0 s4.0 s~equal
time to Green4.0 s4.0 s (at bfb's polling floor)~equal
cpu9.0 s12.4 s1.38x
cpu, % of wall146%187%
waiting for a core4.2 ms1,180.2 ms279.04x
migrations01,910
faults min/maj13,130 / 049,093 / 127,992
peak RSS144.7 MiB761.8 MiB5.27x
disk written6.5 MiB36.9 MiB
disk read010.8 MiB

W0 d=4 floor: graph traversal with the distance taken out 1.20x

closed-loop · collection=bench0 · client -p 1 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second4,8664,0471.20x
client p50201 µs243 µs1.21x
client p99249 µs303 µs1.22x
client p99.9315 µs374 µs1.19x
server p50110 µs135 µs1.24x
server p99121 µs172 µs1.42x
wall clock10.3 s12.4 s1.20x
queries sent50,00050,000
cpu7.1 s12.6 s1.79x
cpu, % of wall68%102%
waiting for a core4.4 ms3.7 ms0.85x
migrations0149
faults min/maj0 / 0282 / 0
peak RSS144.7 MiB760.9 MiB5.26x
disk written00
disk read00
1.20x is 1.65x less work per query x 0.65x cores busy during the row (0.65 against 1.00) x 1.00x clock x 1.12x counted on-CPU share: the middle term is occupancy, not search speed d=4 floor: this measures graph traversal and heaps with the distance taken out, not request plumbing (findings 28: 76% of the server's CPU was HNSW search, 24% kernel; the RPC path is a small remainder)

W1 ingest throughput

upload · collection=bench1 · client -p 8 -t 8 -c 1 (-c bfb default) · n=105,100
metricstrawmANNQdrant
wall clock3.2 s4.7 s1.47x
of which upload1.0 s2.6 s2.48x
cpu2.0 s18.2 s8.92x
cpu, % of wall64%390%
waiting for a core1.0 ms7,564.6 ms7,565.16x
migrations04,367
faults min/maj225,897 / 0160,649 / 251,721
peak RSS1.1 GiB2.5 GiB2.33x
disk written821.1 MiB2.2 GiB
disk read0444.0 KiB

W2 index build time

upload · collection=bench2 · client -p 8 -t 8 -c 1 (-c bfb default) · n=105,100
metricstrawmANNQdrant
wall clock26.2 s27.3 s1.04x
of which upload1.0 s2.4 s2.29x
of which index wait23.1 s23.1 s~equal
time to Green23.1 s23.1 s~equal
cpu167.8 s144.6 s0.86x
cpu, % of wall639%529%
waiting for a core99.0 ms7,150.1 ms72.22x
migrations05,343
faults min/maj229,765 / 0261,652 / 321,681
peak RSS1.1 GiB4.0 GiB3.58x
disk written821.1 MiB3.7 GiB
disk read012.0 MiB

W3 search, fp32, single query 1.30x

ef=128 · closed-loop · collection=bench2 · client -p 1 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,2819861.30x
recall@100.99860.9970~equal
client p50759 µs1.00 ms1.32x
client p991.40 ms1.43 ms1.02x
client p99.91.73 ms1.54 ms0.89x
server p50660 µs887 µs1.34x
server p991.28 ms1.31 ms1.03x
wall clock39.1 s50.8 s1.30x
queries sent50,00050,000
cpu35.6 s51.7 s1.45x
cpu, % of wall91%102%
waiting for a core3.3 ms4.9 ms1.50x
migrations0174
faults min/maj5 / 01,802 / 0
peak RSS1.1 GiB4.0 GiB3.58x
disk written00
disk read00

W4 search, saturating (closed loop) 0.82x

ef=128 · closed-loop · collection=bench2 · client -p 64 -t 16 -c 2 · n=50,000
metricstrawmANNQdrant
queries/second4,6565,6830.82x
recall@100.99860.9970~equal
client p5013.69 ms10.97 ms0.80x
client p9915.28 ms22.55 ms1.48x
client p99.934.24 ms35.02 ms1.02x
server p5013.57 ms10.44 ms0.77x
server p9915.13 ms18.06 ms1.19x
wall clock10.8 s8.8 s0.82x
queries sent50,00050,000
cpu77.3 s69.6 s0.90x
cpu, % of wall718%788%
waiting for a core7.7 ms67,279.5 ms8,776.39x
migrations07,023
faults min/maj84 / 05,372 / 0
peak RSS1.1 GiB4.0 GiB3.58x
disk written00
disk read00

W4-sat50 search, fixed rate at 50% of saturation (open loop) offered

ef=128 · open-loop · offered=2,328/s (50% of saturation, measured 4,656 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second2,3282,841offered
recall@100.99860.9970~equal
client p50889 µs1.12 ms1.26x
client p991.57 ms2.09 ms1.33x
client p99.91.80 ms2.49 ms1.38x
server p50784 µs966 µs1.23x
server p991.46 ms1.91 ms1.31x
wall clock86.0 s70.5 s0.82x
queries sent200,000200,000
cpu167.8 s224.8 s1.34x
cpu, % of wall195%319%
waiting for a core24.3 ms22,057.5 ms907.25x
migrations0497,175
faults min/maj12 / 03,260 / 0
peak RSS1.1 GiB4.0 GiB3.58x
disk written00
disk read00

W4-sat70 search, fixed rate at 70% of saturation (open loop) offered

ef=128 · open-loop · offered=3,259/s (70% of saturation, measured 4,656 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second3,2593,978offered
recall@100.99860.9970~equal
client p501.01 ms1.30 ms1.28x
client p992.12 ms2.65 ms1.25x
client p99.92.47 ms3.56 ms1.44x
server p50903 µs1.11 ms1.23x
server p991.99 ms2.38 ms1.20x
wall clock61.5 s50.4 s0.82x
queries sent200,000200,000
cpu196.9 s244.4 s1.24x
cpu, % of wall320%485%
waiting for a core26.2 ms93,503.0 ms3,571.19x
migrations0583,329
faults min/maj20 / 02,241 / 0
peak RSS1.1 GiB4.0 GiB3.58x
disk written00
disk read00

W4-sat90 search, fixed rate at 90% of saturation (open loop) offered

ef=128 · open-loop · offered=4,191/s (90% of saturation, measured 4,656 qps) · collection=bench2 · client -p n/a under --rps -t 16 -c 2 · n=200,000
metricstrawmANNQdrant
queries/second4,1915,115offered
recall@100.99860.9970~equal
client p501.15 ms1.91 ms1.66x
client p992.57 ms8.26 ms3.21x
client p99.93.51 ms13.36 ms3.80x
server p501.03 ms1.54 ms1.49x
server p992.42 ms6.02 ms2.48x
wall clock47.9 s39.2 s0.82x
queries sent200,000200,000
cpu222.4 s259.1 s1.17x
cpu, % of wall465%660%
waiting for a core29.4 ms161,417.8 ms5,486.89x
migrations0267,514
faults min/maj48 / 02,445 / 0
peak RSS1.1 GiB4.0 GiB3.58x
disk written00
disk read00

W5 search batched (16 distinct dataset queries per request) 2.15x

ef=128 · closed-loop · collection=bench2 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second4,6562,1642.15x
recall@100.99860.9970~equal
client p506.86 ms14.78 ms2.15x
client p998.12 ms16.64 ms2.05x
client p99.98.55 ms17.30 ms2.02x
server p506.26 ms13.98 ms2.24x
server p997.52 ms15.81 ms2.10x
wall clock10.7 s23.1 s2.15x
queries sent50,00050,000
cpu76.3 s45.7 s0.60x
cpu, % of wall710%198%
waiting for a core7.2 ms9.6 ms1.33x
migrations01,280
faults min/maj0 / 01,284 / 0
peak RSS1.1 GiB4.0 GiB3.58x
disk written00
disk read00
per-batch latency (16 queries/request); queries: dataset, random-sample 2.15x is 0.59x less work per query x 3.60x cores busy during the row (7.10 against 1.97) x 1.00x clock x 1.01x counted on-CPU share: the middle term is occupancy, not search speed batched 16, so qps and rps differ by 16x; the ratio uses qps

W6-upload scalar quantization: load

upload · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=105,100
metricstrawmANNQdrant
wall clock88.4 s24.1 s0.27x
of which upload1.0 s2.9 s2.82x
of which index wait85.2 s19.1 s0.22x
time to Green85.2 s19.1 s0.22x
cpu227.9 s85.1 s0.37x
cpu, % of wall258%354%
waiting for a core107.3 ms3,386.9 ms31.55x
migrations33,6191,206.33x
faults min/maj716,129 / 0432,188 / 355,587
peak RSS4.0 GiB6.0 GiB1.51x
disk written821.1 MiB4.6 GiB
disk read014.2 MiB

W6 quantized: scalar 1.33x

ef=128 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second3,6212,7261.33x
recall@100.99890.9974~equal
client p50543 µs728 µs1.34x
client p99842 µs938 µs1.11x
client p99.91.12 ms1.03 ms0.92x
server p50447 µs609 µs1.36x
server p99728 µs813 µs1.12x
wall clock13.8 s18.4 s1.33x
queries sent50,00050,000
cpu24.1 s37.2 s1.54x
cpu, % of wall174%203%
waiting for a core6.3 ms10.9 ms1.71x
migrations014,687
faults min/maj0 / 0985 / 0
peak RSS4.0 GiB6.0 GiB1.51x
disk written00
disk read00

W6-ef32 SQ8 recall control, ef=32 (latency only) 1.45x

ef=32 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=3.2 · rescore=true
metricstrawmANNQdrant
queries/second7,6195,2681.45x
recall@100.99470.9885~equal
client p50240 µs375 µs1.56x
client p99435 µs464 µs1.07x
client p99.9509 µs563 µs1.11x
server p50154 µs263 µs1.71x
server p99325 µs337 µs1.04x
wall clock6.6 s9.5 s1.44x
queries sent50,00050,000
cpu10.2 s19.2 s1.88x
cpu, % of wall155%202%
waiting for a core2.2 ms8.5 ms3.79x
migrations09,455
faults min/maj0 / 0122 / 0
peak RSS4.0 GiB6.0 GiB1.51x
disk written00
disk read00
1.45x is 1.76x less work per query x 0.78x cores busy during the row (1.55 against 2.00) x 1.00x clock x 1.06x counted on-CPU share: the middle term is occupancy, not search speed

W6-ef64 SQ8 recall control, ef=64 (latency only) 1.27x

ef=64 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=6.4 · rescore=true
metricstrawmANNQdrant
queries/second5,0583,9851.27x
recall@100.99750.9948~equal
client p50377 µs498 µs1.32x
client p99700 µs618 µs0.88x
client p99.9793 µs716 µs0.90x
server p50284 µs385 µs1.36x
server p99583 µs495 µs0.85x
wall clock9.9 s12.6 s1.27x
queries sent50,00050,000
cpu16.3 s25.5 s1.57x
cpu, % of wall164%203%
waiting for a core3.8 ms9.1 ms2.38x
migrations011,474
faults min/maj0 / 0129 / 0
peak RSS4.0 GiB6.0 GiB1.51x
disk written00
disk read00

W6-ef128 SQ8 recall control, ef=128 (latency only) 1.32x

ef=128 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second3,6152,7311.32x
recall@100.99890.9974~equal
client p50543 µs726 µs1.34x
client p99860 µs936 µs1.09x
client p99.91.14 ms1.02 ms0.89x
server p50447 µs608 µs1.36x
server p99746 µs811 µs1.09x
wall clock13.9 s18.3 s1.32x
queries sent50,00050,000
cpu24.2 s37.3 s1.54x
cpu, % of wall175%203%
waiting for a core2.9 ms10.8 ms3.75x
migrations014,753
faults min/maj0 / 0476 / 0
peak RSS4.0 GiB6.0 GiB1.51x
disk written00
disk read00

W6-ef256 SQ8 recall control, ef=256 (latency only) 1.38x

ef=256 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=25.6 · rescore=true
metricstrawmANNQdrant
queries/second2,3971,7341.38x
recall@100.99950.9986~equal
client p50826 µs1.14 ms1.39x
client p991.10 ms1.52 ms1.38x
client p99.91.18 ms1.63 ms1.37x
server p50728 µs1.02 ms1.40x
server p99996 µs1.39 ms1.39x
wall clock20.9 s28.9 s1.38x
queries sent50,00050,000
cpu38.2 s58.5 s1.53x
cpu, % of wall183%202%
waiting for a core4.2 ms12.9 ms3.11x
migrations018,679
faults min/maj0 / 0640 / 0
peak RSS4.0 GiB6.0 GiB1.51x
disk written00
disk read00

W6-ef512 SQ8 recall control, ef=512 (latency only) 1.42x

ef=512 · closed-loop · collection=bench6 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=51.2 · rescore=true
metricstrawmANNQdrant
queries/second1,4231,0051.42x
recall@100.99970.9993~equal
client p501.39 ms1.98 ms1.42x
client p991.86 ms2.63 ms1.41x
client p99.92.01 ms2.80 ms1.39x
server p501.29 ms1.84 ms1.43x
server p991.76 ms2.49 ms1.41x
wall clock35.2 s49.8 s1.42x
queries sent50,00050,000
cpu66.7 s100.7 s1.51x
cpu, % of wall190%202%
waiting for a core6.3 ms71.4 ms11.31x
migrations020,776
faults min/maj4 / 01,376 / 0
peak RSS4.0 GiB6.0 GiB1.51x
disk written00
disk read00

W7-upload binary quantization: load

upload · collection=bench7 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=105,100
metricstrawmANNQdrant
wall clock28.3 s16.1 s0.57x
of which upload1.0 s2.9 s2.87x
of which index wait25.1 s11.0 s0.44x
time to Green25.1 s11.0 s0.44x
cpu167.5 s37.8 s0.23x
cpu, % of wall592%235%
waiting for a core502.4 ms2,105.4 ms4.19x
migrations53,933786.60x
faults min/maj237,989 / 0328,498 / 413,120
peak RSS4.0 GiB8.4 GiB2.10x
disk written821.1 MiB4.5 GiB
disk read015.5 MiB

W7 quantized: binary + oversampling parity

ef=128 · closed-loop · collection=bench7 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second4,3524,354parity
recall@100.98270.9851~equal
client p50433 µs455 µs1.05x
client p99677 µs536 µs0.79x
client p99.9841 µs651 µs0.77x
server p50338 µs342 µs~equal
server p99563 µs401 µs0.71x
wall clock11.5 s11.5 s~equal
queries sent50,00050,000
cpu19.5 s23.1 s1.18x
cpu, % of wall169%201%
waiting for a core5.8 ms12.1 ms2.09x
migrations08,960
faults min/maj0 / 0506 / 0
peak RSS4.0 GiB8.4 GiB2.10x
disk written00
disk read00
within the ±2.6% band this dataset's noise floor puts on W7: no measured difference, not a small one

W8-upload PQ: load

upload · collection=bench8 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=105,100
metricstrawmANNQdrant
wall clock176.7 s150.2 s0.85x
of which upload1.0 s2.8 s2.71x
of which index wait173.6 s145.4 s0.84x
time to Green173.6 s145.4 s0.84x
cpu1,362.4 s955.9 s0.70x
cpu, % of wall771%636%
waiting for a core433.9 ms2,287.5 ms5.27x
migrations115,66415,664.00x
faults min/maj461,956 / 0252,009 / 287,843
peak RSS5.4 GiB8.4 GiB1.57x
disk written821.1 MiB3.6 GiB
disk read8.0 KiB13.5 MiB

W8 quantized: PQ 1.30x

ef=128 · closed-loop · collection=bench8 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000 · oversampling=12.8 · rescore=true
metricstrawmANNQdrant
queries/second1,4521,1211.30x
recall@100.99820.9963~equal
client p501.36 ms1.76 ms1.30x
client p991.93 ms2.26 ms1.17x
client p99.92.28 ms2.57 ms1.13x
server p501.26 ms1.62 ms1.29x
server p991.83 ms2.08 ms1.14x
wall clock34.5 s44.6 s1.29x
queries sent50,00050,000
cpu65.3 s89.9 s1.38x
cpu, % of wall189%202%
waiting for a core4.9 ms136.5 ms27.67x
migrations020,325
faults min/maj0 / 05,327 / 0
peak RSS5.4 GiB8.4 GiB1.57x
disk written00
disk read00

W9 exact / brute force parity

exact · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=2,000
metricstrawmANNQdrant
queries/second7171parity
client p50109.30 ms113.19 ms1.04x
client p99164.37 ms128.73 ms0.78x
client p99.9180.92 ms140.29 ms0.78x
server p50108.48 ms111.83 ms1.03x
server p99163.69 ms127.43 ms0.78x
wall clock28.1 s28.1 s~equal
queries sent2,0002,000
cpu197.4 s221.5 s1.12x
cpu, % of wall702%788%
waiting for a core28.8 ms11,013.9 ms382.93x
migrations08,603
faults min/maj0 / 0787 / 0
peak RSS5.4 GiB8.4 GiB1.57x
disk written00
disk read00
within the ±3.1% band this dataset's noise floor puts on W9: no measured difference, not a small one brute force over the whole collection, no index involved

W10-ef32 recall control, ef=32 (latency only) 1.15x

ef=32 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second11,96710,4281.15x
recall@100.99430.9877~equal
client p50658 µs714 µs1.08x
client p99983 µs1.39 ms1.42x
client p99.91.22 ms1.80 ms1.48x
server p50556 µs540 µs0.97x
server p99874 µs1.16 ms1.33x
wall clock4.2 s4.8 s1.15x
queries sent50,00050,000
cpu28.9 s27.9 s0.96x
cpu, % of wall686%576%
waiting for a core4.6 ms10,904.3 ms2,364.29x
migrations072,162
faults min/maj0 / 01,315 / 0
peak RSS5.4 GiB8.4 GiB1.57x
disk written00
disk read00
1.15x is 0.92x less work per query x 1.21x cores busy during the row (6.96 against 5.77) x 1.00x clock x 1.04x counted on-CPU share: the middle term is occupancy, not search speed

W10-ef64 recall control, ef=64 (latency only) 1.05x

ef=64 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second7,4767,1411.05x
recall@100.99720.9945~equal
client p501.05 ms1.03 ms~equal
client p991.64 ms2.06 ms1.26x
client p99.91.87 ms2.63 ms1.40x
server p50946 µs841 µs0.89x
server p991.53 ms1.79 ms1.17x
wall clock6.7 s7.0 s1.05x
queries sent50,00050,000
cpu47.1 s43.5 s0.92x
cpu, % of wall701%617%
waiting for a core5.6 ms28,791.1 ms5,173.36x
migrations0129,916
faults min/maj0 / 0973 / 0
peak RSS5.4 GiB8.4 GiB1.57x
disk written00
disk read00

W10-ef128 recall control, ef=128 (latency only) 0.97x

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second4,6364,7910.97x
recall@100.99860.9970~equal
client p501.69 ms1.56 ms0.92x
client p992.76 ms3.01 ms1.09x
client p99.93.19 ms3.83 ms1.20x
server p501.58 ms1.32 ms0.83x
server p992.65 ms2.69 ms~equal
wall clock10.8 s10.5 s0.97x
queries sent50,00050,000
cpu76.5 s66.2 s0.86x
cpu, % of wall707%632%
waiting for a core9.1 ms41,856.1 ms4,600.43x
migrations0150,106
faults min/maj4 / 0577 / 0
peak RSS5.4 GiB8.4 GiB1.57x
disk written00
disk read00

W10-ef256 recall control, ef=256 (latency only) 0.91x

ef=256 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,8523,1200.91x
recall@100.99920.9985~equal
client p502.75 ms2.46 ms0.89x
client p994.63 ms4.57 ms~equal
client p99.95.37 ms5.64 ms1.05x
server p502.63 ms2.19 ms0.83x
server p994.51 ms4.17 ms0.92x
wall clock17.6 s16.1 s0.91x
queries sent50,00050,000
cpu124.4 s105.4 s0.85x
cpu, % of wall708%656%
waiting for a core13.4 ms58,014.1 ms4,315.22x
migrations0178,287
faults min/maj4 / 0724 / 0
peak RSS5.4 GiB8.4 GiB1.57x
disk written00
disk read00

W10-ef512 recall control, ef=512 (latency only) 0.88x

ef=512 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,7311,9590.88x
recall@100.99970.9993~equal
client p504.52 ms3.99 ms0.88x
client p997.75 ms7.07 ms0.91x
client p99.98.93 ms8.50 ms0.95x
server p504.40 ms3.69 ms0.84x
server p997.63 ms6.63 ms0.87x
wall clock28.9 s25.6 s0.88x
queries sent50,00050,000
cpu204.1 s174.0 s0.85x
cpu, % of wall706%681%
waiting for a core22.7 ms80,763.0 ms3,558.50x
migrations0197,452
faults min/maj0 / 01,495 / 0
peak RSS5.4 GiB8.4 GiB1.57x
disk written00
disk read00

W12-upload filtered search: load 105,100 with payloads

upload · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=105,100
metricstrawmANNQdrant
wall clock26.3 s50.3 s1.91x
of which upload1.1 s3.7 s3.53x
of which index wait23.1 s44.1 s1.91x
time to Green23.1 s44.1 s1.91x
cpu166.8 s219.2 s1.31x
cpu, % of wall634%436%
waiting for a core94.9 ms7,688.9 ms81.06x
migrations14,1854,185.00x
faults min/maj230,557 / 0214,691 / 352,826
peak RSS5.6 GiB8.6 GiB1.54x
disk written821.1 MiB3.7 GiB
disk read12.0 KiB24.6 MiB

W12-sel1 filtered search, one keyword (~1% of bench12) parity

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,5942,503parity
recall@101.00000.9999~equal
client p50767 µs797 µs1.04x
client p99864 µs949 µs1.10x
client p99.9930 µs1.03 ms1.10x
server p50670 µs675 µs~equal
server p99762 µs809 µs1.06x
wall clock19.3 s20.0 s1.04x
queries sent50,00050,000
cpu34.8 s40.6 s1.17x
cpu, % of wall180%203%
waiting for a core3.0 ms10.9 ms3.64x
migrations013,621
faults min/maj0 / 0673 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00
within the ±9.1% band this dataset's noise floor puts on W12-sel1: no measured difference, not a small one filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set

W12-sel10 filtered search, any of 10 keywords (~10% of bench12) parity

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,1471,280parity
recall@100.99910.9952~equal
client p501.74 ms1.57 ms0.90x
client p992.48 ms1.96 ms0.79x
client p99.92.68 ms2.08 ms0.78x
server p501.64 ms1.43 ms0.87x
server p992.38 ms1.81 ms0.76x
wall clock43.6 s39.1 s0.90x
queries sent50,00050,000
cpu83.3 s79.1 s0.95x
cpu, % of wall191%202%
waiting for a core4.9 ms48.9 ms10.07x
migrations020,685
faults min/maj0 / 0712 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00
within the ±12.7% band this dataset's noise floor puts on W12-sel10: no measured difference, not a small one filtered search over a keyword payload index both engines were asked to build; whether each did is read back from the engine and refused beside the row when it did not. Recall is measured under the filter, against ground truth restricted to the matching set

W12-sel1-ef32 filtered recall control, one keyword, ef=32 (latency only) 0.57x

ef=32 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,6054,5990.57x
recall@101.00000.9974~equal
client p50764 µs433 µs0.57x
client p99861 µs529 µs0.61x
client p99.9970 µs642 µs0.66x
server p50668 µs318 µs0.48x
server p99758 µs395 µs0.52x
wall clock19.2 s10.9 s0.57x
queries sent50,00050,000
cpu34.7 s22.1 s0.64x
cpu, % of wall180%202%
waiting for a core3.2 ms8.8 ms2.76x
migrations010,523
faults min/maj0 / 0204 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00

W12-sel1-ef64 filtered recall control, one keyword, ef=64 (latency only) 0.75x

ef=64 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,5993,4860.75x
recall@101.00000.9996~equal
client p50765 µs574 µs0.75x
client p99859 µs692 µs0.80x
client p99.9934 µs795 µs0.85x
server p50668 µs455 µs0.68x
server p99757 µs555 µs0.73x
wall clock19.3 s14.4 s0.75x
queries sent50,00050,000
cpu34.7 s29.2 s0.84x
cpu, % of wall180%203%
waiting for a core5.1 ms12.2 ms2.40x
migrations012,016
faults min/maj0 / 0289 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00

W12-sel1-ef128 filtered recall control, one keyword, ef=128 (latency only) parity

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,6122,510parity
recall@101.00000.9999~equal
client p50762 µs795 µs1.04x
client p99848 µs944 µs1.11x
client p99.9939 µs1.02 ms1.09x
server p50665 µs673 µs~equal
server p99743 µs806 µs1.08x
wall clock19.2 s20.0 s1.04x
queries sent50,00050,000
cpu34.6 s40.4 s1.17x
cpu, % of wall180%203%
waiting for a core2.1 ms13.2 ms6.29x
migrations014,175
faults min/maj0 / 0292 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00
within the ±9.1% band this dataset's noise floor puts on W12-sel1-ef128: no measured difference, not a small one

W12-sel1-ef256 filtered recall control, one keyword, ef=256 (latency only) 1.49x

ef=256 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,5931,7421.49x
recall@101.00001.0000~equal
client p50766 µs1.14 ms1.49x
client p99863 µs1.32 ms1.53x
client p99.9936 µs1.40 ms1.50x
server p50669 µs1.02 ms1.52x
server p99756 µs1.18 ms1.56x
wall clock19.3 s28.7 s1.49x
queries sent50,00050,000
cpu34.7 s58.1 s1.67x
cpu, % of wall179%202%
waiting for a core1.9 ms20.9 ms10.85x
migrations014,920
faults min/maj0 / 0593 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00

W12-sel1-ef512 filtered recall control, one keyword, ef=512 (latency only) 2.29x

ef=512 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,6061,1372.29x
recall@101.00001.0000~equal
client p50763 µs1.76 ms2.30x
client p99857 µs1.95 ms2.27x
client p99.9936 µs2.07 ms2.21x
server p50666 µs1.62 ms2.43x
server p99751 µs1.79 ms2.38x
wall clock19.2 s44.0 s2.29x
queries sent50,00050,000
cpu34.6 s88.2 s2.55x
cpu, % of wall180%200%
waiting for a core5.1 ms117.3 ms23.01x
migrations013,954
faults min/maj0 / 0942 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00

W12-sel10-ef32 filtered recall control, any of 10 keywords, ef=32 (latency only)

ef=32 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,5832,998-
recall@100.99290.8703
client p50767 µs661 µs
client p991.05 ms862 µs
client p99.91.24 ms965 µs
server p50670 µs537 µs
server p99952 µs726 µs
wall clock19.4 s16.7 s
queries sent50,00050,000
cpu35.2 s33.8 s
cpu, % of wall182%202%
waiting for a core9.3 ms10.3 ms
migrations014,208
faults min/maj0 / 0538 / 0
peak RSS5.6 GiB8.6 GiB
disk written00
disk read00
recall unequal: 0.9929 vs 0.8703; §7.4 compares at equal recall

W12-sel10-ef64 filtered recall control, any of 10 keywords, ef=64 (latency only)

ef=64 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,7612,081-
recall@100.99750.9666
client p501.13 ms963 µs
client p991.59 ms1.20 ms
client p99.91.75 ms1.31 ms
server p501.03 ms834 µs
server p991.49 ms1.06 ms
wall clock28.4 s24.1 s
queries sent50,00050,000
cpu53.1 s48.7 s
cpu, % of wall187%202%
waiting for a core3.9 ms16.8 ms
migrations016,801
faults min/maj0 / 0266 / 0
peak RSS5.6 GiB8.6 GiB
disk written00
disk read00
recall unequal: 0.9975 vs 0.9666; §7.4 compares at equal recall

W12-sel10-ef128 filtered recall control, any of 10 keywords, ef=128 (latency only) parity

ef=128 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second1,1461,268parity
recall@100.99910.9952~equal
client p501.75 ms1.59 ms0.91x
client p992.49 ms1.97 ms0.79x
client p99.92.70 ms2.09 ms0.77x
server p501.64 ms1.45 ms0.88x
server p992.38 ms1.81 ms0.76x
wall clock43.7 s39.5 s0.90x
queries sent50,00050,000
cpu83.5 s79.8 s0.96x
cpu, % of wall191%202%
waiting for a core5.6 ms62.9 ms11.16x
migrations020,657
faults min/maj0 / 0940 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00
within the ±12.7% band this dataset's noise floor puts on W12-sel10-ef128: no measured difference, not a small one

W12-sel10-ef256 filtered recall control, any of 10 keywords, ef=256 (latency only) 0.52x

ef=256 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second4087840.52x
recall@101.00000.9992~equal
client p504.88 ms2.59 ms0.53x
client p995.61 ms3.18 ms0.57x
client p99.95.90 ms3.34 ms0.57x
server p504.56 ms2.43 ms0.53x
server p994.98 ms3.02 ms0.61x
wall clock122.6 s63.8 s0.52x
queries sent50,00050,000
cpu231.5 s128.6 s0.56x
cpu, % of wall189%202%
waiting for a core31.8 ms76.6 ms2.41x
migrations019,993
faults min/maj0 / 01,458 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00

W12-sel10-ef512 filtered recall control, any of 10 keywords, ef=512 (latency only) 0.85x

ef=512 · closed-loop · collection=bench12 · client -p 2 -t 2 -c 1 (-p, -t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second4144860.85x
recall@101.00000.9999~equal
client p504.81 ms4.17 ms0.87x
client p995.34 ms5.16 ms0.97x
client p99.95.67 ms5.41 ms0.95x
server p504.55 ms3.96 ms0.87x
server p994.99 ms4.92 ms~equal
wall clock120.7 s102.8 s0.85x
queries sent50,00050,000
cpu231.3 s204.8 s0.89x
cpu, % of wall192%199%
waiting for a core35.0 ms87.0 ms2.49x
migrations022,174
faults min/maj0 / 02,574 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00

W13 scroll / pagination 1.09x

closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=200,000
metricstrawmANNQdrant
queries/second50,84546,4501.09x
client p50138 µs158 µs1.15x
client p99232 µs304 µs1.31x
client p99.9276 µs466 µs1.69x
server p5011 µs28 µs2.59x
server p9916 µs95 µs5.85x
wall clock4.1 s2.2 s0.55x
queries sent200,000100,000
cpu5.6 s9.2 s1.65x
cpu, % of wall137%414%
waiting for a core2.4 ms932.5 ms396.38x
migrations0112,908
faults min/maj0 / 0425 / 0
peak RSS5.6 GiB8.6 GiB1.54x
disk written00
disk read00
strawmANN: -n 200000 (4x the table's 50000) Qdrant: -n 100000 (2x the table's 50000) strawmANN: 0.13 ms of the client's 0.14 ms p50 is not server time (13x), so 92% of this row is the load generator and the socket Qdrant: 0.13 ms of the client's 0.16 ms p50 is not server time (6x), so 82% of this row is the load generator and the socket 1.09x is 3.29x less work per query x 0.33x cores busy during the row (1.39 against 4.24) x 1.00x clock x 1.02x counted on-CPU share: the middle term is occupancy, not search speed

W11-steady mixed read/write below the rebuild threshold: search bench2 while 5,255 synthetic points append

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second2,3883,857-
client p503.10 ms1.83 ms
client p996.75 ms5.33 ms
client p99.97.80 ms16.33 ms
server p502.98 ms1.52 ms
server p996.62 ms4.66 ms
wall clock21.0 s13.0 s
queries sent50,00050,000
cpu147.1 s96.9 s
cpu, % of wall701%745%
waiting for a core20.7 ms77,335.9 ms
migrations0171,847
faults min/maj12,338 / 0403,201 / 130,786
peak RSS5.6 GiB8.6 GiB
disk written41.2 MiB5.9 GiB
disk read038.3 MiB
strawmANN: search covered 79% of the append; append 198 points/s Qdrant: search covered 49% of the append; append 198 points/s search-during-write; no recall join

W11 mixed read/write: search bench2 while 21,020 synthetic points append (runs last)

ef=128 · closed-loop · collection=bench2 · client -p 8 -t 2 -c 1 (-t, -c bfb default) · n=50,000
metricstrawmANNQdrant
queries/second6233,453-
client p5010.83 ms2.07 ms
client p9927.03 ms6.82 ms
client p99.931.01 ms17.67 ms
server p5010.64 ms1.69 ms
server p9926.34 ms5.84 ms
wall clock80.3 s14.5 s
queries sent50,00050,000
cpu606.7 s156.9 s
cpu, % of wall756%1,081%
waiting for a core309,613.4 ms100,964.5 ms
migrations71182,017
faults min/maj58,385 / 01,670,641 / 591,145
peak RSS5.6 GiB23.8 GiB
disk written164.5 MiB25.7 GiB
disk read0152.9 MiB
strawmANN: write overlap 88%; append 299 points/s Qdrant: search covered 21% of the append; append 299 points/s search-during-write; no recall join; the writer covered under 90% of strawmANN's search, so the last 12% of that row measured the rebuild the append provoked rather than a concurrent write
Host discipline the §7.1 gate, per check, and ambient load per row

Ambient load per row

The §7.1 gate checks the machine once, at the start. Colour is the engine, as everywhere else; a hatched bar with a red edge is a row that had another process on the box while it ran, which the gate cannot see.

AMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical) gate pass

One machine: both runs were measured on it, and the checks below are the same for both. What differed at each run start is listed underneath, per run.
okAMD Ryzen AI 9 HX PRO 370 w/ Radeon 890M (24 logical)
okgovernor=performance
okboost disabled
oksmt=on
declared rather than required (§7.1 asks for SMT on or off, not
for off), and hashed, so these rows never share a chart with
SMT-off ones
cpu N and cpu N+12 share a core (12 pairs);
a --server-cpus or --client-cpus set holding both cpus of a
pair measures the engine on half the cores it names
okprofile=as-deployed (isolcpus=none nohz_full=none)
the scheduler is running as it ships, so these numbers describe a
deployment rather than the engine in isolation; they may not be
compared against an `isolated` run, and the environment hash
enforces that
okthp=madvise numa_nodes=1 numa_balancing=0
okperf_event_paranoid=-1
oktask_delayacct=1, so time blocked on a device is measurable
okperf at /usr/bin/perf, so `workloads.py run --perf` can attach
okcgroup io controller reaches the engines' own scope, so block-layer read/write operations are measurable
okzig=0.16.0
busy: Xorg(6%)
measurement profile: as-deployed
strawmANN at run start · environment hash 373e8531fd3d1bf6
Qdrant at run start · environment hash 373e8531fd3d1bf6
What this does not establish

No Qdrant developer has used this tool. Nothing here has been run, checked against something already known, or disagreed with by anyone outside the project: every number has one author and one reviewer, and they are the same person. A result you can refute is more useful to us than one you accept.

It has never been pointed at a real Qdrant regression. The harness detects a known ISA slowdown and correctly reports no difference on a control row. That is internal consistency, not evidence it would flag a regression in your tree or stay quiet through a refactor. bench/harness/qdrant_ab.py exists to run two of your commits through it blind, and that experiment has not been done.

Part of the measured throughput gap is kernel width, not architecture. Read from a dev checkout: Qdrant's distance path tops out at four 256-bit accumulators and has no AVX-512, while these kernels use 512-bit ones on a machine that has them. That is a real difference and it is not the same claim as "the design is faster".

The engines are not equivalent, by construction. §1 removes sharding, replication, consensus, snapshots, sparse vectors, multivectors and disk-resident operation — permanently. Payload storage and filtering were in that list until 2026-09-03 and are now built (M7), so W12 measures a filtered search with a keyword index on both engines. Anything still unbuilt answers UNIMPLEMENTED naming the construct and reports n/a, never a silent degradation. A strawman that was not faster would mean it was badly built; the question is by how much, and where the model was wrong.

Reproducing this

Everything below runs from a clone. The engine has no dependencies; the analysis path is a uv project pinned by bench/uv.lock.

scripts/doctor.py                    # what can this machine measure?
bench/setup.py check                 # §7.1 host gate: governor, boost, isolation, idle
conformance/datasets/datasets.py fetch

# one engine at a time, §7.1
zig build -Doptimize=ReleaseFast
# --connections matters: W4 opens ~32 sockets (-t 16 -c 2) and a server with
# fewer closes the excess before the HTTP/2 preface, which the client reports
# only as "transport error".
./zig-out/bin/strawmann --port 6334 --connections 64
bench/harness/workloads.py run http://localhost:6334 strawmann --sink --report

# the other half of every number: recall, and the licence to compare at all
cd conformance
cargo run --release -- relevance --engine http://localhost:6334 --label strawmann \
  --base $DATA/sift1m.fbin --queries $DATA/sift1m_query.fbin \
  --ground-truth $DATA/gt/sift1m.euclid.k100.gt.json --metric euclid \
  --ef 32 64 128 256 512 --json ../bench/results/strawmann/recall.json
cargo run --release -- differ --strawmann http://localhost:6334 --qdrant http://localhost:6434 \
  --base $DATA/sift1m.fbin --queries $DATA/sift1m_query.fbin \
  --ground-truth $DATA/gt/sift1m.euclid.k100.gt.json \
  --json ../bench/results/strawmann/conformance.json

uv run --project bench bench/harness/report.py strawmann qdrant

Disagreements are the point. docs/spec.md is what all of this is measured against, docs/bugs.md lists the measurement bugs found so far (each produced a plausible wrong number rather than a failure), and docs/validation.md records what has and has not been checked.