Production retrieval for long documents
Fast enough to sit inside a product your users wait on. Every number below says when it was measured and how.
Seconds, not minutes.
Timed server-side on 3M 2018 Form 10-K (160 pages): question in, sections or a cited answer out.
From question to the sections that hold the answer.
Median 1.76 s · p90 2.61 s · 12 runs, 8 Oct 2026
The full answer, written, with every pin cite attached.
Median 3.46 s · p90 3.98 s · 12 runs, 8 Oct 2026
From upload to the first question it can answer.
13 s for the same filing · measured 7 Oct 2026
Method: Live on the production API, timed server-side. 12 runs: six questions, each asked twice, against the same filing. On nine other long documents measured the same day (regulations, a trial protocol, a patent manual), a full answer took a median of 6 s; we are working that down. Measured 8 Oct 2026. Upload to ready timed on the same filing, 7 Oct 2026.
It finds the passage that holds the answer.
19 real 10-K filings, 40 questions, run 21 Sep 2026, against chunk-and-embed and BM25 on the same questions.
Answer in the top five
92.5%
The passage holding the answer was in the first five results for 37 of 40 questions.
- Vectorless92.5% · 37 of 40
- Chunk-and-embed37.5% · 15 of 40
- BM2520%
Answer in the first result
75%
The very first result held the answer span.
- Vectorless75%
- Chunk-and-embed22.5%
- BM257.5%
Method: Each question is graded on retrieval: whether the passage that holds the answer was among the results. The two baselines are chunk-and-embed, using BGE-small embeddings, and BM25 keyword ranking. Run 21 Sep 2026. Benchmark harness: vectorless-bench.
It reads the document the way it was written.
No chunking
The document keeps the structure its authors gave it: contents, parts, sections. Nothing is cut into fixed-size pieces and embedded.
It opens what a careful reader would
An LLM reads the outline and opens the sections that bear on the question, the way an analyst works through a filing.
Every answer can be replayed
Answers carry pin cites (section, page range, exact quote) and a trace of every section opened, so anyone can check how it got there.
- Form 10-K (opened)
- Part I
- Item 1. Business
- Item 1A. Risk factors
- Part II (opened)
- Item 7. Management's discussion and analysis (opened)
- Results of operations
- Liquidity and capital resources (cited)
- Item 8. Financial statements (opened)
- Item 7. Management's discussion and analysis (opened)
- Part IV
- Part I
- 01Read the contents
- 02Opened Part II
- 03Opened Item 7
- 04Opened Item 8
- 05Cited Liquidity and capital resources
- Section
- Item 7 › Liquidity and capital resources
- Pages
- The page range it sits on
- Quote
- “The exact sentence the answer rests on.”
Run it where your data already lives.
The hosted API
Upload a document, ask a question, get a cited answer. Nothing to run.
One binary on Postgres
Run the engine inside your own network. One binary, one Postgres database, no vector store to operate.
Whichever LLM you already use
Point it at the model provider you already have a contract with. Your documents go where your model already goes.
Send us one document and the question your system gets wrong.
Or put it through the free plan yourself and read the pin cites and trace it returns.