Hybrid Retrieval Benchmark — dbpedia

500 sampled queries, seed 42, warmup 10, concurrency 8. Dense: all-MiniLM-L6-v2. Sparse: Qdrant/bm25. Colbert: answerai-colbert-small-v1. Cross-encoder: local fastembed rerank (Xenova/ms-marco-MiniLM-L-6-v2). Dataset: BeIR/dbpedia-entity-generated-queries, 100K passages, one generated query per passage.
recall@20 by quantization config × rescorer (hybrid prefetch, prefetch_limit=100)
configcolbertcross-encoderrrf
no-quant (f32)96.496.095.0
no-quant (f16)96.095.895.0
no-quant (turbo4)96.296.095.2
scalar int896.496.095.0
product x1696.496.095.2
binary (f32)96.095.895.2
binary (f16)95.696.095.0
turbo4 tq196.495.895.0
recall / mrr / ndcg by rescorer
recall@20 mrr@20 ndcg@20 (opacity = metric; hue = rescorer)
0.00 0.25 0.50 0.75 1.00 0.962 0.884 0.903 colbert 0.959 0.886 0.904 cross-encoder 0.951 0.808 0.843 rrf score (0-1) rescorer (hybrid prefetch, k=20, prefetch_limit=100)
throughput by rescorer
0.5 1 5 10 50 100 300 66.35 qps colbert 0.43 qps cross-encoder 285.33 qps rrf throughput, qps (log) rescorer, concurrency=8 (hybrid, k=20, prefetch_limit=100)
recall@20 vs prefetch_limit, by rescorer (hybrid, k=20)
colbert · recall@20
0.94 0.95 0.97 pf=25 pf=50 pf=100 96.2%
cross-encoder · recall@20
0.94 0.95 0.97 pf=25 pf=50 pf=100 95.9%
rrf · recall@20
0.94 0.95 0.97 pf=25 pf=50 pf=100 95.1%
throughput vs prefetch_limit, by rescorer (hybrid, k=20)
colbert · qps (log)
1 10 100 pf=25 pf=50 pf=100 66.35
cross-encoder · qps (log)
1 10 100 pf=25 pf=50 pf=100 0.43
rrf · qps (log)
1 10 100 pf=25 pf=50 pf=100 285.33
colbert & cross-encoder, hybrid vs sparse-only prefetch (k=20, prefetch_limit=100)
recall@20 mrr@20 ndcg@20 (opacity = metric; hue = rescorer)
hybrid prefetch
0.00 0.25 0.50 0.75 1.00 0.962 colbert 0.959 cross-encoder
colbertmrr 0.884ndcg 0.903
cross-encodermrr 0.886ndcg 0.904
sparse-only prefetch
0.00 0.25 0.50 0.75 1.00 0.908 colbert 0.906 cross-encoder
colbertmrr 0.853ndcg 0.867
cross-encodermrr 0.856ndcg 0.869
recall@20 vs prefetch_limit, colbert & cross-encoder (sparse-only, k=20)
colbert · recall@20
0.86 0.89 0.92 pf=25 pf=50 pf=100 90.8%
cross-encoder · recall@20
0.86 0.89 0.92 pf=25 pf=50 pf=100 90.6%
throughput vs prefetch_limit, colbert & cross-encoder (sparse-only, k=20)
colbert · qps (log)
1 10 100 pf=25 pf=50 pf=100 75.04
cross-encoder · qps (log)
1 10 100 pf=25 pf=50 pf=100 0.51
mrr@20 by config, colbert rescorer (bold outline = turbo4 datatype)
hybridsparse-only
0.680 0.753 0.827 0.900 no-quant (f32) no-quant (f16) no-quant (turbo4) scalar int8 product x16 binary (f32) binary (f16) turbo4 tq1
ndcg@20 by config, colbert rescorer (bold outline = turbo4 datatype)
hybridsparse-only
0.740 0.800 0.860 0.920 no-quant (f32) no-quant (f16) no-quant (turbo4) scalar int8 product x16 binary (f32) binary (f16) turbo4 tq1
exact values, colbert rescorer, k=20 prefetch_limit=100
confighybrid mrrhybrid ndcgsparse-only mrrsparse-only ndcg
no-quant (f32)0.88600.90460.85400.8673
no-quant (f16)0.88570.90370.85400.8673
no-quant (turbo4)0.88400.90270.85460.8676
scalar int80.88390.90310.85200.8657
product x160.87960.89980.85010.8643
binary (f32)0.88470.90280.85320.8665
binary (f16)0.88390.90130.85400.8669
turbo4 tq10.88410.90330.85460.8677