Advanced search controls
Metadata filters, semantic ratio tuning, boosts and chunk relations, all through one query API. The lever you need is a request parameter, not a fork of the pipeline.
Now self-tuningcaptain_eval
The complete file search platform. Synced with cloud storage, fast and accurate, and optimized for complex retrieval.
Powerful ingestion and self-tuning retrieval,
behind one query API, ready for production scale.
Sub-second queries through a deterministic API, with filters, rerank and candidate depth all yours to configure.
Sub-second queries through a deterministic API, with filters, rerank and candidate depth all yours to configure.
Sync files and self-improve search as one platform
Every layer above ships together, on one index. Nothing to wire and nothing to keep aligned.
claude mcp add --transport http captain https://mcp.captain.dev/mcp \ --header "Authorization: Bearer cap_..."
A DOCX or a scanned PDF is not text yet. Whatever the parser drops, no search can find.
XML fragments in a zip, with tables that only exist as formatting. Captain's parser was tuned on the ugliest DOCX files our customers could find.
Clean PDFs parse anywhere. Captain also reads the rotated scan and the handwritten margin note.
Transcribed, described and indexed at ingest. The chart on slide 12 is a query, not a scrubbing session.
A public benchmark tells you how a system ran on someone else's corpus. Run and Captain builds hundreds of Q&A evals from your own files, scores retrieval against them, and tunes until it hits your target.
When evaluated against other leading file-search APIs on production queries, alternatives fail to answer about one in three questions.
See what people search for, and exactly where answers fall short. The feedback loop makes retrieval better over time.
See the questions your data has yet to answer, surfaced automatically, so you know exactly where to close the gaps.
Catch black-hole vector attacks (arXiv:2604.05480) before they reach production, a failure mode most DIY search layers never see.
Metadata filters, semantic ratio tuning, boosts and chunk relations, all through one query API. The lever you need is a request parameter, not a fork of the pipeline.
One index across text, documents and video. Ask where the Q3 margin chart appears and get the file, the page and the timestamp.
Index straight from object storage and cloud drives. One platform across every source, not a collection per silo.
Redact sensitive customer data at index time with one flag, before it reaches your models or logs.

The fully-managed file search API.
Point it at your files and search them. Parsing, embeddings, vector DB and evaluation, all managed for you.
The main interface for Captain.
Call captain_eval from the hosted MCP server and it scores retrieval on evals written from your own files, then tunes toward the accuracy or latency target you set.

Bucket change notifications drive the index in near real time, and scheduled reconciliation jobs sweep for anything a notification dropped.
Your buckets remain the source of truth.
TLS in transit, encryption at rest, and opt-in PII masking so sensitive data is hidden from your agents.