Evaluations.

Retrieval benchmarks with the dataset, scoring and comparison systems named on each report.

BEIR text retrieval benchmark

September 2026

Captain averaged 0.647 nDCG@10 on 1,271 queries over 66,454 documents. Cloudflare AI Search averaged 0.451, Gemini File Search 0.432 and Elasticsearch v8 0.386.

BEIR nDCG@10 and share of queries returning results for Captain, Elasticsearch v8, Gemini File Search and Cloudflare AI Search, overall and across SciFact, NFCorpus and FiQA.
nDCG@10CaptainElasticsearchv8GeminiFile SearchCloudflareAI Search
Overall0.6470.3860.4320.451
SciFact0.8870.6100.6850.681
NFCorpus0.4400.2950.2860.295
FiQA0.6160.2540.3250.378
Queries returning results
SciFact100%100%81%58 of 300 return nothing100%
NFCorpus100%100%90%32 of 323 return nothing100%
FiQA100%100%69%201 of 648 return nothing100%
Indexing speed
Time to index each corpus for Captain, Elasticsearch v8, Gemini File Search and Cloudflare AI Search.
Time to index the corpus
SciFact (5,183 documents)3 min10 s27 min73 min
NFCorpus (3,633 documents)2 min7 s18 min59 min
FiQA (57,638 documents)20 min60 s4.85 hrs40.1 hrs

Captain runs at semantic ratio 1.0 with no reranker, the setting a sweep on 400 held-out queries picked; at its out-of-the-box defaults it averages 0.555. Cloudflare AI Search is at its defaults on the same corpora and queries. Elasticsearch v8 and Gemini File Search are the April 2026 run.

MRAG multimodal retrieval benchmarkApril 2026
02MRAG-BenchApril 2026

MRAG multimodal
retrieval benchmark.

Captain on the MRAG-Bench multimodal retrieval evaluation from ICLR 2025: 1,251 questions over 16,130 corpus images.

SystemOverall accuracy
  • Captain81.3%
  • GPT-4o + CLIP RAG68.96%
  • Gemini Pro + CLIP RAG65.93%
  • Claude 3.5 + CLIP RAG63.56%
  • Human + Retrieved RAG61.38%
  • Needle in a Haystack subset100%

Evaluation code and methodology: github.com/runcaptain/captain-mrag-bench

MRAG-Bench evaluation results showing Captain at 81.3% overall accuracy
Captain Open RAG BenchmarkMarch 2026

Captain scored 95% retrieval accuracy on the Open RAG Benchmark with an LLM as judge, against 81% for a reranker RAG pipeline and 78% for vanilla RAG.

  • Captain95%
  • Reranker RAG81%
  • Vanilla RAG78%
Head to head: accuracy, win rate, wins, losses and ties per pipeline over 595 to 599 pairwise games judged by an LLM.
Head to headAccuracyWin rateWinsLossesTies
Captain94.6%86.7%5154930
Reranker RAG80.5%19.0%113326156
Vanilla RAG78.5%15.9%95348156
Bradley-Terry leaderboard, Elo-scaled rating per pipeline.
#Bradley-TerryRating
1Captain16571.570
2Reranker RAG1431-0.695
3Vanilla RAG1412-0.875
Method
  • Four indexing strategies combined: multimodal embeddings, contextualized embeddings, BM25, full-text search.
  • Embeddings normalised so semantic representations stay consistent across modalities.
  • Chunk boundaries chosen with cosine-distance conditionals.