Document Retrieval activates only when the question contract and available context show a concrete evidence gap that local documents can resolve.
Document Retrieval
Make private and oversized corpora answerable with provenance intact.
Consumes canonical structured evidence and combines lexical, optional semantic and conservative graph routes with bounded evidence packing, citation provenance and explicit diagnostics.
A bounded evidence packet containing the most relevant passages, their structural and source provenance, retrieval diagnostics and enough context for Answer Assurance to verify later claims.
Why it exists
The problem this module is designed to solve.
Large document collections exceed useful context windows, while similarity-only retrieval can miss exact identifiers, tables, chronology and cross-document relationships. Making retrieval mandatory also degrades simple cases that would be better handled directly.
02 / How it works
A bounded path from need to accountable result.
The public model below describes responsibilities and decisions, not sensitive implementation details, provider secrets or customer data.
- 01
Preserve document structure
Turn headings, paragraphs, tables, pages and source relationships into canonical evidence rather than flattening everything into anonymous text.
- 02
Choose the retrieval route
Use exact lexical, structured, optional semantic or conservative relationship paths according to the request instead of forcing one index strategy.
- 03
Build a bounded evidence packet
Balance relevance, source diversity, requested facets and context budget while retaining the passages needed to verify the result.
- 04
Return provenance and diagnostics
Expose which route contributed each item, what could not be found and whether retrieval is sufficient for the requested answer.
Retrieval engine / public technical brief
From heterogeneous evidence to a bounded, citable packet.
Document Retrieval is not a generic vector-search wrapper. It keeps a deterministic path, adds semantic or relationship routes selectively, and preserves the source structure required to inspect every returned item.
Evidence pipeline
Six explicit stages, with provenance carried through every one.
Changing an index, embedding model or reranker must not require reinterpreting the original source. The canonical evidence layer separates extraction from retrieval representation.
- 01
Ingestion handoff
Accept authenticated canonical evidence from the media-specific route. The retrieval engine never assumes that a filename extension proves content or extraction quality.
- 02
Canonicalize
Keep text, tables, figures, images, audio spans and selected frames as typed evidence with hierarchy, page, bounding-box, time, speaker and provider provenance.
- 03
Represent and index
Build deterministic lexical records first, then add dense, learned-sparse or structured representations only when the configured profile and evidence justify them.
- 04
Retrieve under scope
Apply tenant and workspace filters on every query, plus optional conversation and source allow-lists, before any candidate can enter the evidence set.
- 05
Fuse and bound
Preserve raw references while combining eligible routes. Reranking is limited to a bounded candidate set and is promoted only after it beats the simpler baseline.
- 06
Pack and cite
Return a context-budgeted packet with source locators, route diagnostics, scores, warnings and exact references for downstream answer assurance.
Retrieval strategies
Exact lookup stays first-class. Semantic search remains optional.
The engine can compose several signals, but no route is entitled to run merely because it exists. Exact identifiers, structured fields and short facts reserve a deterministic lexical path.
Lexical and BM25
Deterministic retrieval for identifiers, clauses, names, rare terms and quoted text. It is model-free, inspectable and remains the fallback when optional routes fail.
Dense semantic
Local or private embeddings can recover paraphrases and conceptual similarity. They remain profile-bound, replaceable and disabled when the measured gain does not justify memory or latency.
Learned sparse
A separate sparse leg can expand vocabulary while retaining token-level signals. It is not presented as the lexical baseline or silently merged into dense search.
Structured filters
Tables, metadata, dates, facets and access attributes remain queryable as fields instead of becoming an untraceable text blob.
Root and hierarchy
Parent-child paths reconnect a matching passage to its document root, section and neighboring evidence while keeping document boundaries intact.
Bounded evidence graph
Support, typed-relation, adjacency, sequence and shared-concept edges can fill unused evidence slots for neighborhood or multi-hop questions. Traversal starts only from baseline candidates, cannot rewrite their scores, has strict bounds and falls back to the unchanged baseline. It requires neither a graph database nor a graph model.
Source families
The engine sees typed evidence, not just extracted text.
Format authentication, conversion, OCR, ASR and visual interpretation belong to Multimodal Ingestion and its licensed adapters. Retrieval consumes only the evidence those routes produced and qualified.
Native text and data
Parse structure directly; keep exact terms and fields available to lexical and structured retrieval.
Documents and publications
Preserve headings, lists, sheets, slides, tables, pages and embedded media instead of flattening the source.
Scans and images
Consume OCR, layout, visual regions and metadata only when the configured ingestion adapter produced them.
Audio and conversations
Index transcript spans on the original clock; attach anonymous speaker spans only when diarization evidence exists.
Video and timed media
Join ASR spans, scene boundaries, keyframes and selected frame OCR without erasing their independent timestamps.
Email, archives and specialist files
Retain container and parent-child provenance; unsupported semantics remain opaque rather than being guessed.
Recognition is intentionally broader than qualified extraction. The complete generated registry, common image list, truth ladder and real-file evidence live on the ingestion brief.
Execution and failure
Immediate when sufficient. Resumable or lazy when evidence is expensive.
Complex sources are checkpointed by source checksum, version, stage and provider. A failed enrichment step does not erase verified text, metadata or an already valid index representation.
Bypass retrieval
If authorized context already answers the request, the system can use it directly. Retrieval is not made mandatory for simple cases.
Return the sufficient path
Small native sources and deterministic lexical queries can complete in the request path when their bounded budget permits it.
Resume by stage
Large parsing, indexing and media workflows can checkpoint work and continue without restarting completed stages.
Repair only critical gaps
Low-quality OCR, sparse transcripts or missing optional evidence can be queued for later recovery while usable sibling evidence remains available.
Keep media clocks independent
Video can run audio and selected-frame routes in parallel, then reconcile evidence by source time rather than by processing order.
Preserve verified evidence
Unavailable embeddings, rerankers, graph expansion or enrichers degrade to a simpler route with warnings instead of invalidating the whole source.
Measured evidence
Promotion follows comparative results, including the failures.
These results describe frozen internal profiles. They are evidence for the named implementation and test corpus, not production service levels, customer-corpus relevance or certification.
Root-to-intent recall at five
Zero wrong-document top-one results, zero recorded scope leaks and 2.091 ms local release-build p95 on one sealed 48-query synthetic BM25 holdout.
Adapter-owned scope; one development host; no customer or production generalization.Top-one on the visible diagnostic holdout
Multilingual E5 qint8 reached 29/30 top-three versus 11/30 lexical top-one, with 4 ms release p95 on Apple M1 and 970 MB standalone RSS.
The tested BM25+dense reciprocal-rank fusion fell to 21/30 top-one and was not promoted.Baseline-preserving expansion
The first-party graph contract defines scoped paths, budgets, diagnostics and failure fallback without a graph database or model dependency.
Adaptive graph routing and broad quality improvement remain promotion gates.Operational control
Replaceable infrastructure, stable evidence identity.
- OpenSearch is the default external search adapter; Qdrant remains a challenger, not a lock-in.
- OCR, ASR, vision, embedding and reranking providers stay behind explicit licensed adapters.
- Traces record route, provider, stage, version, timing, counts, scores and references without raw bodies, credentials or embeddings.
- Every record and query carries tenant and workspace scope; optional source and conversation allow-lists narrow it further.
03 / Customer and operator value
Improves coverage on large or multi-document tasks while allowing direct full-context inference to win when it is simpler and better.
Structure-aware evidence packets
Lexical, semantic and graph-ready signals
Transient compact source index from authorized structural metadata
Exact source-root and scope filtering
Verified Qdrant and Milvus read paths
Atomic local index recovery
Exact source provenance
Digest-bound LongMemEval, LoCoMo, ConvoMem and MemBench adapters
Aligned public-artifact retrieval leads with comparison limits stated
Fail-closed retrieval snapshot validation
Opt-in local hybrid recall with an already-installed Ollama embedding model
Atomic vector rebuilds that preserve source records and resume unchanged vectors
Scoped recall filters over caller-managed Memory locations and recorded-time windows
04 / Where it creates value
Concrete situations, not generic feature claims.
These are representative product situations. Every deployment still requires its own policy, data boundary and acceptance criteria.
Contracts and policy libraries
Recover exact clauses, identifiers, exceptions and linked provisions without losing the source hierarchy.
Multi-document investigations
Connect events, decisions and supporting records across reports while keeping each document's provenance separate.
Large technical knowledge bases
Combine exact lookup with semantic discovery for terminology, procedures, tables and rare operational facts.
05 / Role in the cognitive system
A clear responsibility creates a trustworthy boundary.
No module is allowed to become an invisible monolith. It owns a narrow contract, composes with named capabilities and refuses responsibilities that belong elsewhere.
What it owns
- Authorized source catalogue
- Compact structural source index
- Indexable evidence
- Search and ranking
- Retrieval diagnostics
- Evidence packing
- Index recovery
- Provenance
What it composes with
What it refuses to own
Boundary before convenience.
It does not generate final answers, choose an LLM or persist user preferences.
06 / Vision and mission
Evidence when direct context is not enough
Retrieval serves the mission by adding complexity only for a demonstrated gap, then returning evidence in a form that remains inspectable and accountable through publication.
AI systems that remain useful, inspectable and sovereign across models, providers and deployment boundaries.
Build the cognitive layer that chooses the smallest sufficient path and turns evidence into accountable action.
Capture the value of advanced AI without surrendering data control, architectural freedom or intellectual honesty.
07 / Evidence and maturity
What the current label means — and what remains open.
Passed with moderate confidence
Typed retrieval, scope isolation and supported kernel/server composition are implemented and locally verified; the evidence does not establish production-scale relevance or latency.
- No production-scale asynchronous ingestion campaign.
- No representative customer-corpus drift, relevance or latency qualification.
- No independent assurance opinion.
What exists today
Structure-aware retrieval, bounded evidence packing and provenance contracts are implemented. The Retrieval crate owns the transient source index exposed as `mentaview_retrieval_source_index` and authenticated `/retrieval/source-index`; it returns metadata and caller-query terms only. The Retrieval page now carries the LongMemEval, LoCoMo, ConvoMem and MemBench comparisons, Qdrant and Milvus assurance, and local index recovery evidence; none is counted as Governed Memory capability. A sealed synthetic 48-query root-to-intent holdout recovered 48/48 expected items with zero wrong-document top-one results and zero scope leaks.
Controlled pilot
A bounded runnable surface exists and can be evaluated in a controlled engagement. Production-scale controls and operating evidence remain gated.
What must earn promotion
Qualify production-scale index construction, representative corpus drift, wider prospective holdouts and customer-specific relevance, recovery and latency targets.
Public truth boundary: M4− is a non-standard Mentaview engineering label for internal laboratory validation. It is not an official TRL decision and does not assert production qualification, customer acceptance, independent assurance, certification or universal performance. Inspect the complete assessment record.