Hybrid Retrieval Benchmark — scifact

100 sampled queries, seed 42, warmup 10, concurrency 8. Dense: all-MiniLM-L6-v2. Sparse: Qdrant/bm25. Colbert: answerai-colbert-small-v1. Cross-encoder: local fastembed rerank (Xenova/ms-marco-MiniLM-L-6-v2). Dataset: BeIR/scifact-generated-queries, 15.4K passages, one generated query per passage.
recall@20 by quantization config × rescorer (hybrid prefetch, prefetch_limit=100)
configcolbertcross-encoderrrf
no-quant (f32)93.092.089.0
no-quant (f16)93.092.089.0
no-quant (turbo4)92.091.090.0
scalar int893.092.089.0
product x1692.093.088.0
binary (f32)92.093.088.0
binary (f16)92.094.089.0
turbo4 tq193.092.089.0
recall / mrr / ndcg by rescorer
recall@20 mrr@20 ndcg@20 (opacity = metric; hue = rescorer)
0.00 0.25 0.50 0.75 1.00 0.925 0.470 0.581 colbert 0.924 0.471 0.580 cross-encoder 0.889 0.432 0.541 rrf score (0-1) rescorer (hybrid prefetch, k=20, prefetch_limit=100)
throughput by rescorer
0.1 0.5 1 5 10 50 100 300 58.59 qps colbert 0.12 qps cross-encoder 272.19 qps rrf throughput, qps (log) rescorer, concurrency=8 (hybrid, k=20, prefetch_limit=100)
recall@20 vs prefetch_limit, by rescorer (hybrid, k=20)
colbert · recall@20
0.87 0.90 0.93 pf=25 pf=50 pf=100 92.5%
cross-encoder · recall@20
0.87 0.90 0.93 pf=25 pf=50 pf=100 92.4%
rrf · recall@20
0.87 0.90 0.93 pf=25 pf=50 pf=100 88.9%
throughput vs prefetch_limit, by rescorer (hybrid, k=20)
colbert · qps (log)
1 10 100 pf=25 pf=50 pf=100 58.59
cross-encoder · qps (log)
1 10 100 pf=25 pf=50 pf=100 0.12
rrf · qps (log)
1 10 100 pf=25 pf=50 pf=100 272.19
colbert & cross-encoder, hybrid vs sparse-only prefetch (k=20, prefetch_limit=100)
recall@20 mrr@20 ndcg@20 (opacity = metric; hue = rescorer)
hybrid prefetch
0.00 0.25 0.50 0.75 1.00 0.925 colbert 0.924 cross-encoder
colbertmrr 0.470ndcg 0.581
cross-encodermrr 0.471ndcg 0.580
sparse-only prefetch
0.00 0.25 0.50 0.75 1.00 0.899 colbert 0.909 cross-encoder
colbertmrr 0.478ndcg 0.582
cross-encodermrr 0.469ndcg 0.575
recall@20 vs prefetch_limit, colbert & cross-encoder (sparse-only, k=20)
colbert · recall@20
0.84 0.88 0.92 pf=25 pf=50 pf=100 89.9%
cross-encoder · recall@20
0.84 0.88 0.92 pf=25 pf=50 pf=100 90.9%
throughput vs prefetch_limit, colbert & cross-encoder (sparse-only, k=20)
colbert · qps (log)
1 10 100 pf=25 pf=50 pf=100 64.94
cross-encoder · qps (log)
1 10 100 pf=25 pf=50 pf=100 0.12
mrr@20 by config, colbert rescorer (bold outline = turbo4 datatype)
hybridsparse-only
0.380 0.433 0.487 0.540 no-quant (f32) no-quant (f16) no-quant (turbo4) scalar int8 product x16 binary (f32) binary (f16) turbo4 tq1
ndcg@20 by config, colbert rescorer (bold outline = turbo4 datatype)
hybridsparse-only
0.520 0.557 0.593 0.630 no-quant (f32) no-quant (f16) no-quant (turbo4) scalar int8 product x16 binary (f32) binary (f16) turbo4 tq1
exact values, colbert rescorer, k=20 prefetch_limit=100
confighybrid mrrhybrid ndcgsparse-only mrrsparse-only ndcg
no-quant (f32)0.41260.53870.47110.5764
no-quant (f16)0.52270.62150.48510.5873
no-quant (turbo4)0.48760.59230.47530.5785
scalar int80.48600.59480.48150.5839
product x160.43360.55140.45380.5633
binary (f32)0.45580.56870.48440.5866
binary (f16)0.50130.60300.49130.5894
turbo4 tq10.46270.57650.48590.5869