Proof · measured and dated

Production retrieval for long documents

Fast enough to sit inside a product your users wait on. Every number below says when it was measured and how.

Latency · live on the production API

Seconds, not minutes.

Timed server-side on 3M 2018 Form 10-K (160 pages): question in, sections or a cited answer out.

Retrieval~2 s

From question to the sections that hold the answer.

Median 1.76 s · p90 2.61 s · 12 runs, 8 Oct 2026

Cited answer3–4 s

The full answer, written, with every pin cite attached.

Median 3.46 s · p90 3.98 s · 12 runs, 8 Oct 2026

Upload to ready~15 s

From upload to the first question it can answer.

13 s for the same filing · measured 7 Oct 2026

Method: Live on the production API, timed server-side. 12 runs: six questions, each asked twice, against the same filing. On nine other long documents measured the same day (regulations, a trial protocol, a patent manual), a full answer took a median of 6 s; we are working that down. Measured 8 Oct 2026. Upload to ready timed on the same filing, 7 Oct 2026.

Retrieval · FinanceBench

It finds the passage that holds the answer.

19 real 10-K filings, 40 questions, run 21 Sep 2026, against chunk-and-embed and BM25 on the same questions.

Answer in the top five

92.5%

The passage holding the answer was in the first five results for 37 of 40 questions.

  • Vectorless92.5% · 37 of 40
  • Chunk-and-embed37.5% · 15 of 40
  • BM2520%

Answer in the first result

75%

The very first result held the answer span.

  • Vectorless75%
  • Chunk-and-embed22.5%
  • BM257.5%

Method: Each question is graded on retrieval: whether the passage that holds the answer was among the results. The two baselines are chunk-and-embed, using BGE-small embeddings, and BM25 keyword ranking. Run 21 Sep 2026. Benchmark harness: vectorless-bench.

How it finds things

It reads the document the way it was written.

01

No chunking

The document keeps the structure its authors gave it: contents, parts, sections. Nothing is cut into fixed-size pieces and embedded.

02

It opens what a careful reader would

An LLM reads the outline and opens the sections that bear on the question, the way an analyst works through a filing.

03

Every answer can be replayed

Answers carry pin cites (section, page range, exact quote) and a trace of every section opened, so anyone can check how it got there.

The filing's own structureIllustrative
  • Form 10-K (opened)
    • Part I
      • Item 1. Business
      • Item 1A. Risk factors
    • Part II (opened)
      • Item 7. Management's discussion and analysis (opened)
        • Results of operations
        • Liquidity and capital resources (cited)
      • Item 8. Financial statements (opened)
    • Part IV
OpenedCitedNot needed
Trace
  1. 01Read the contents
  2. 02Opened Part II
  3. 03Opened Item 7
  4. 04Opened Item 8
  5. 05Cited Liquidity and capital resources
Pin cite
Section
Item 7 › Liquidity and capital resources
Pages
The page range it sits on
Quote
“The exact sentence the answer rests on.”
Illustrative: a filing's outline with the sections opened to answer a question, the trace of each step, and the pin cite the answer carries.
In production

Run it where your data already lives.

Hosted

The hosted API

Upload a document, ask a question, get a cited answer. Nothing to run.

Self-hosted

One binary on Postgres

Run the engine inside your own network. One binary, one Postgres database, no vector store to operate.

Your model

Whichever LLM you already use

Point it at the model provider you already have a contract with. Your documents go where your model already goes.

Try it on yours

Send us one document and the question your system gets wrong.

Or put it through the free plan yourself and read the pin cites and trace it returns.