Hybrid Retrieval Benchmark — msmarco

500 sampled queries, seed 42, warmup 10, concurrency 8. Dense: all-MiniLM-L6-v2. Sparse: Qdrant/bm25. Colbert: answerai-colbert-small-v1. Cross-encoder: local fastembed rerank (Xenova/ms-marco-MiniLM-L-6-v2). Dataset: jordane95/msmarco-passage-corpus-with-query, 100K passages, real associated search queries (no synthetic/LLM-generated queries).
recall@20 by quantization config × rescorer (hybrid prefetch, prefetch_limit=100)
configcolbertcross-encoderrrf
no-quant (f32)96.897.495.2
no-quant (f16)96.897.495.2
no-quant (turbo4)96.697.495.4
scalar int897.097.495.2
product x1696.697.695.6
binary (f32)96.697.495.2
binary (f16)96.897.495.2
turbo4 tq196.497.295.4
recall / mrr / ndcg by rescorer
recall@20 mrr@20 ndcg@20 (opacity = metric; hue = rescorer)
0.00 0.25 0.50 0.75 1.00 0.967 0.710 0.772 colbert 0.974 0.717 0.779 cross-encoder 0.953 0.645 0.719 rrf score (0-1) rescorer (hybrid prefetch, k=20, prefetch_limit=100)
throughput by rescorer
0.5 1 5 10 50 100 300 60.93 qps colbert 0.41 qps cross-encoder 256.89 qps rrf throughput, qps (log) rescorer, concurrency=8 (hybrid, k=20, prefetch_limit=100)
recall@20 vs prefetch_limit, by rescorer (hybrid, k=20)
colbert · recall@20
0.94 0.96 0.98 pf=25 pf=50 pf=100 96.7%
cross-encoder · recall@20
0.94 0.96 0.98 pf=25 pf=50 pf=100 97.4%
rrf · recall@20
0.94 0.96 0.98 pf=25 pf=50 pf=100 95.3%
throughput vs prefetch_limit, by rescorer (hybrid, k=20)
colbert · qps (log)
1 10 100 pf=25 pf=50 pf=100 60.93
cross-encoder · qps (log)
1 10 100 pf=25 pf=50 pf=100 0.41
rrf · qps (log)
1 10 100 pf=25 pf=50 pf=100 256.89
colbert & cross-encoder, hybrid vs sparse-only prefetch (k=20, prefetch_limit=100)
recall@20 mrr@20 ndcg@20 (opacity = metric; hue = rescorer)
hybrid prefetch
0.00 0.25 0.50 0.75 1.00 0.967 colbert 0.974 cross-encoder
colbertmrr 0.710ndcg 0.772
cross-encodermrr 0.717ndcg 0.779
sparse-only prefetch
0.00 0.25 0.50 0.75 1.00 0.926 colbert 0.930 cross-encoder
colbertmrr 0.699ndcg 0.754
cross-encodermrr 0.701ndcg 0.757
recall@20 vs prefetch_limit, colbert & cross-encoder (sparse-only, k=20)
colbert · recall@20
0.86 0.90 0.94 pf=25 pf=50 pf=100 92.7%
cross-encoder · recall@20
0.86 0.90 0.94 pf=25 pf=50 pf=100 93.0%
throughput vs prefetch_limit, colbert & cross-encoder (sparse-only, k=20)
colbert · qps (log)
1 10 100 pf=25 pf=50 pf=100 68.26
cross-encoder · qps (log)
1 10 100 pf=25 pf=50 pf=100 0.42
mrr@20 by config, colbert rescorer (bold outline = turbo4 datatype)
hybridsparse-only
0.650 0.673 0.697 0.720 no-quant (f32) no-quant (f16) no-quant (turbo4) scalar int8 product x16 binary (f32) binary (f16) turbo4 tq1
ndcg@20 by config, colbert rescorer (bold outline = turbo4 datatype)
hybridsparse-only
0.720 0.740 0.760 0.780 no-quant (f32) no-quant (f16) no-quant (turbo4) scalar int8 product x16 binary (f32) binary (f16) turbo4 tq1
exact values, colbert rescorer, k=20 prefetch_limit=100
confighybrid mrrhybrid ndcgsparse-only mrrsparse-only ndcg
no-quant (f32)0.71510.77630.70060.7558
no-quant (f16)0.71480.77600.69950.7550
no-quant (turbo4)0.70560.76880.69250.7491
scalar int80.70970.77270.70690.7607
product x160.69870.76300.70000.7549
binary (f32)0.71500.77580.70030.7552
binary (f16)0.71520.77630.69950.7551
turbo4 tq10.70500.76790.69260.7492