Module 05 / Retrieve structure, not just similar text

M4 internal labControlled pilot

Document Retrieval

Make private and oversized corpora answerable with provenance intact.

Consumes canonical structured evidence and combines lexical, optional semantic and conservative graph routes with bounded evidence packing, citation provenance and explicit diagnostics.

ACTIVATES WHEN

Document Retrieval activates only when the question contract and available context show a concrete evidence gap that local documents can resolve.

RETURNS

A bounded evidence packet containing the most relevant passages, their structural and source provenance, retrieval diagnostics and enough context for Answer Assurance to verify later claims.

Why it exists

The problem this module is designed to solve.

Large document collections exceed useful context windows, while similarity-only retrieval can miss exact identifiers, tables, chronology and cross-document relationships. Making retrieval mandatory also degrades simple cases that would be better handled directly.

02 / How it works

A bounded path from need to accountable result.

The public model below describes responsibilities and decisions, not sensitive implementation details, provider secrets or customer data.

  1. 01

    Preserve document structure

    Turn headings, paragraphs, tables, pages and source relationships into canonical evidence rather than flattening everything into anonymous text.

  2. 02

    Choose the retrieval route

    Use exact lexical, structured, optional semantic or conservative relationship paths according to the request instead of forcing one index strategy.

  3. 03

    Build a bounded evidence packet

    Balance relevance, source diversity, requested facets and context budget while retaining the passages needed to verify the result.

  4. 04

    Return provenance and diagnostics

    Expose which route contributed each item, what could not be found and whether retrieval is sufficient for the requested answer.

Retrieval engine / public technical brief

From heterogeneous evidence to a bounded, citable packet.

Document Retrieval is not a generic vector-search wrapper. It keeps a deterministic path, adds semantic or relationship routes selectively, and preserves the source structure required to inspect every returned item.

4independent retrieval legs
0mandatory model calls on the lexical path
48/48root-to-intent recall at five on the sealed synthetic holdout
1canonical evidence contract across media

Evidence pipeline

Six explicit stages, with provenance carried through every one.

Changing an index, embedding model or reranker must not require reinterpreting the original source. The canonical evidence layer separates extraction from retrieval representation.

  1. 01

    Ingestion handoff

    Accept authenticated canonical evidence from the media-specific route. The retrieval engine never assumes that a filename extension proves content or extraction quality.

  2. 02

    Canonicalize

    Keep text, tables, figures, images, audio spans and selected frames as typed evidence with hierarchy, page, bounding-box, time, speaker and provider provenance.

  3. 03

    Represent and index

    Build deterministic lexical records first, then add dense, learned-sparse or structured representations only when the configured profile and evidence justify them.

  4. 04

    Retrieve under scope

    Apply tenant and workspace filters on every query, plus optional conversation and source allow-lists, before any candidate can enter the evidence set.

  5. 05

    Fuse and bound

    Preserve raw references while combining eligible routes. Reranking is limited to a bounded candidate set and is promoted only after it beats the simpler baseline.

  6. 06

    Pack and cite

    Return a context-budgeted packet with source locators, route diagnostics, scores, warnings and exact references for downstream answer assurance.

Retrieval strategies

Exact lookup stays first-class. Semantic search remains optional.

The engine can compose several signals, but no route is entitled to run merely because it exists. Exact identifiers, structured fields and short facts reserve a deterministic lexical path.

Implemented baseline

Lexical and BM25

Deterministic retrieval for identifiers, clauses, names, rare terms and quoted text. It is model-free, inspectable and remains the fallback when optional routes fail.

Optional canary

Dense semantic

Local or private embeddings can recover paraphrases and conceptual similarity. They remain profile-bound, replaceable and disabled when the measured gain does not justify memory or latency.

Adapter-ready

Learned sparse

A separate sparse leg can expand vocabulary while retaining token-level signals. It is not presented as the lexical baseline or silently merged into dense search.

Contract implemented

Structured filters

Tables, metadata, dates, facets and access attributes remain queryable as fields instead of becoming an untraceable text blob.

Selector implemented

Root and hierarchy

Parent-child paths reconnect a matching passage to its document root, section and neighboring evidence while keeping document boundaries intact.

Selective first-party contract

Bounded evidence graph

Support, typed-relation, adjacency, sequence and shared-concept edges can fill unused evidence slots for neighborhood or multi-hop questions. Traversal starts only from baseline candidates, cannot rewrite their scores, has strict bounds and falls back to the unchanged baseline. It requires neither a graph database nor a graph model.

Source families

The engine sees typed evidence, not just extracted text.

Format authentication, conversion, OCR, ASR and visual interpretation belong to Multimodal Ingestion and its licensed adapters. Retrieval consumes only the evidence those routes produced and qualified.

TXT, Markdown, CSV, TSV, JSON, YAML and XML

Native text and data

Parse structure directly; keep exact terms and fields available to lexical and structured retrieval.

PDF, DOCX, PPTX, XLSX, HTML, RTF and EPUB

Documents and publications

Preserve headings, lists, sheets, slides, tables, pages and embedded media instead of flattening the source.

JPEG, PNG, GIF, WebP, TIFF, HEIC/HEIF, AVIF and specialist raster families

Scans and images

Consume OCR, layout, visual regions and metadata only when the configured ingestion adapter produced them.

WAV, FLAC, M4A/AAC, Ogg/Opus, MP3 and other authenticated audio

Audio and conversations

Index transcript spans on the original clock; attach anonymous speaker spans only when diarization evidence exists.

MP4, MOV, MKV, WebM and other authenticated containers

Video and timed media

Join ASR spans, scene boundaries, keyframes and selected frame OCR without erasing their independent timestamps.

EML, MSG, MBOX, ZIP, TAR and registered technical families

Email, archives and specialist files

Retain container and parent-child provenance; unsupported semantics remain opaque rather than being guessed.

735 registered identifiers

Recognition is intentionally broader than qualified extraction. The complete generated registry, common image list, truth ladder and real-file evidence live on the ingestion brief.

Open every format and quality route

Execution and failure

Immediate when sufficient. Resumable or lazy when evidence is expensive.

Complex sources are checkpointed by source checksum, version, stage and provider. A failed enrichment step does not erase verified text, metadata or an already valid index representation.

DIRECT

Bypass retrieval

If authorized context already answers the request, the system can use it directly. Retrieval is not made mandatory for simple cases.

SYNCHRONOUS

Return the sufficient path

Small native sources and deterministic lexical queries can complete in the request path when their bounded budget permits it.

ASYNCHRONOUS

Resume by stage

Large parsing, indexing and media workflows can checkpoint work and continue without restarting completed stages.

LAZY BACKGROUND

Repair only critical gaps

Low-quality OCR, sparse transcripts or missing optional evidence can be queued for later recovery while usable sibling evidence remains available.

PARALLEL

Keep media clocks independent

Video can run audio and selected-frame routes in parallel, then reconcile evidence by source time rather than by processing order.

PARTIAL FAILURE

Preserve verified evidence

Unavailable embeddings, rerankers, graph expansion or enrichers degrade to a simpler route with warnings instead of invalidating the whole source.

Measured evidence

Promotion follows comparative results, including the failures.

These results describe frozen internal profiles. They are evidence for the named implementation and test corpus, not production service levels, customer-corpus relevance or certification.

CONTROLLED PILOT · LEXICAL48/48

Root-to-intent recall at five

Zero wrong-document top-one results, zero recorded scope leaks and 2.091 ms local release-build p95 on one sealed 48-query synthetic BM25 holdout.

Adapter-owned scope; one development host; no customer or production generalization.
CANARY · EMBEDDED DENSE28/30

Top-one on the visible diagnostic holdout

Multilingual E5 qint8 reached 29/30 top-three versus 11/30 lexical top-one, with 4 ms release p95 on Apple M1 and 970 MB standalone RSS.

The tested BM25+dense reciprocal-rank fusion fell to 21/30 top-one and was not promoted.
CONTRACT · EVIDENCE GRAPHBounded

Baseline-preserving expansion

The first-party graph contract defines scoped paths, budgets, diagnostics and failure fallback without a graph database or model dependency.

Adaptive graph routing and broad quality improvement remain promotion gates.

Operational control

Replaceable infrastructure, stable evidence identity.

  • OpenSearch is the default external search adapter; Qdrant remains a challenger, not a lock-in.
  • OCR, ASR, vision, embedding and reranking providers stay behind explicit licensed adapters.
  • Traces record route, provider, stage, version, timing, counts, scores and references without raw bodies, credentials or embeddings.
  • Every record and query carries tenant and workspace scope; optional source and conversation allow-lists narrow it further.

Retrieval / MemPalace comparison evidence

Retrieval scores and index operations belong to Retrieval.

These benchmark, backend-assurance and index-recovery results measure evidence selection and retrieval infrastructure. They are deliberately excluded from the Governed Memory score.

Reproducible retrieval check

One hit ahead on the aligned public split.

On the exact published 450-case LongMemEval-S split, Mentaview’s later contextual-hybrid profile found a labelled evidence session in the first five results 444 times.

Mentaview contextual hybrid · recall@5
98.67%
444 / 450 · aligned public split
MemPalace Hybrid v4 · recall@5
98.44%
443 / 450 · artifact count · no LLM
Measured difference
+0.22 pp
one additional retrieved case
Mentaview input identity
SHA-256
d6f21e…3a442

Four-protocol retrieval audit

Four aligned artifact leads, with explicit limits.

Mentaview has digest-bound native adapters for the same LongMemEval, LoCoMo, ConvoMem and MemBench retrieval families published by MemPalace. Each score below states the exact comparison boundary rather than implying global product parity.

LongMemEval · R@5
98.67%
MemPalace artifact 98.44% · +0.22 pp
LoCoMo strict light-stem · R@10
91.96%
MemPalace artifact 88.91% · +3.05 pp · same speaker-text
LoCoMo temporal · category 2
91.23%
MemPalace artifact 90.76% · 321 questions
LoCoMo open-domain · category 3
69.80%
MemPalace artifact 69.96% · 96 questions · remaining gap
LoCoMo enriched light-stem · R@10
93.31%
MemPalace aggregate claim 92.40% · +0.91 pp · richer content
ConvoMem product profile v2 · R@10
95.33%
MemPalace artifact 92.87% · +2.47 pp
MemBench intent-routed · R@5
91.36%
MemPalace artifact 80.33% · +11.04 pp · no model
MemBench aggregative
99.60%
MemPalace 99.30% · +0.30 pp
MemBench knowledge update
97.80%
MemPalace 96.00% · +1.80 pp
MemBench simple facts
96.90%
MemPalace 95.90% · +1.00 pp

All-replica Qdrant reads

Hybrid search plus safe, receipt-bound inspection.

Mentaview now applies provider-native all-replica consistency to four typed Qdrant reads: dense/sparse/RRF search, metadata-only inventory, exact scoped count and restricted facets. Provider output is checked independently before release.

Read consistency
All
mandatory on search · scroll · count · facet
Response scope check
Exact
tenant · workspace · conversation · source · citation
Inventory
≤ 256
metadata only · no body · no vectors
Safe facets
3
source · modality · conversation · exact counts
Hybrid fusion
RRF 60
bounded dense + sparse candidates
Audit receipt
SHA-256
request · scope · raw response · released result

Verified Milvus read/write challenger

Verified schema and writes; scoped strong reads.

The optional Milvus library adapter creates or re-reads the exact collection schema, publishes bounded dense evidence only after Milvus acknowledges every stable ID, prevents broad deletion and keeps search at mandatory Strong consistency. It is not a selectable CLI or server retrieval backend; an embedding application must explicitly supply its transport and scope.

Schema re-read
Exact
fields · dimension · COSINE · dynamic disabled
Upsert acknowledgement
Complete
count + every stable 64-hex ID
Delete authority
Explicit
non-empty scoped source allowlist
Read consistency
Strong
mandatory · not caller-downgradable
Output allowlist
8 fields
no stored vector · no free metadata
Audit proof
SHA-256
request · response · validated result

Two-path local index recovery

Restore lexical service or rebuild every vector in one verified commit.

Mentaview can clear only corrupt embedding metadata for an immediate BM25 fallback, or reconstruct every current knowledge vector with one explicitly selected, already-installed local model. Both operations preview without provider calls or mutation.

Recovery paths
2
targeted clear · complete rebuild
Preview provider calls
0
plan and receipt only
Authorized fields
2
embedding model · vector
Full-plan bound
4,096
all embeddings complete before mutation
Apply confirmation
SHA-256
exact store and reconstruction plan
Remote repair surfaces
0
not exposed through MCP or HTTP

03 / Customer and operator value

Improves coverage on large or multi-document tasks while allowing direct full-context inference to win when it is simpler and better.

01

Structure-aware evidence packets

02

Lexical, semantic and graph-ready signals

03

Transient compact source index from authorized structural metadata

04

Exact source-root and scope filtering

05

Verified Qdrant and Milvus read paths

06

Atomic local index recovery

07

Exact source provenance

08

Digest-bound LongMemEval, LoCoMo, ConvoMem and MemBench adapters

09

Aligned public-artifact retrieval leads with comparison limits stated

10

Fail-closed retrieval snapshot validation

11

Opt-in local hybrid recall with an already-installed Ollama embedding model

12

Atomic vector rebuilds that preserve source records and resume unchanged vectors

13

Scoped recall filters over caller-managed Memory locations and recorded-time windows

04 / Where it creates value

Concrete situations, not generic feature claims.

These are representative product situations. Every deployment still requires its own policy, data boundary and acceptance criteria.

USE CASE 01

Contracts and policy libraries

Recover exact clauses, identifiers, exceptions and linked provisions without losing the source hierarchy.

USE CASE 02

Multi-document investigations

Connect events, decisions and supporting records across reports while keeping each document's provenance separate.

USE CASE 03

Large technical knowledge bases

Combine exact lookup with semantic discovery for terminology, procedures, tables and rare operational facts.

05 / Role in the cognitive system

A clear responsibility creates a trustworthy boundary.

No module is allowed to become an invisible monolith. It owns a narrow contract, composes with named capabilities and refuses responsibilities that belong elsewhere.

What it owns

  • Authorized source catalogue
  • Compact structural source index
  • Indexable evidence
  • Search and ranking
  • Retrieval diagnostics
  • Evidence packing
  • Index recovery
  • Provenance

What it composes with

IngestionQuestion ContractAnswer AssuranceExternal Research

What it refuses to own

Boundary before convenience.

It does not generate final answers, choose an LLM or persist user preferences.

06 / Vision and mission

Evidence when direct context is not enough

Retrieval serves the mission by adding complexity only for a demonstrated gap, then returning evidence in a form that remains inspectable and accountable through publication.

VISION

AI systems that remain useful, inspectable and sovereign across models, providers and deployment boundaries.

MISSION

Build the cognitive layer that chooses the smallest sufficient path and turns evidence into accountable action.

PURPOSE

Capture the value of advanced AI without surrendering data control, architectural freedom or intellectual honesty.

07 / Evidence and maturity

What the current label means — and what remains open.

M4 internal lab
ENGINEERING MATURITYM4Mentaview internal laboratory assessment
OBJECTIVE REVIEW

Passed with moderate confidence

Typed retrieval, scope isolation and supported kernel/server composition are implemented and locally verified; the evidence does not establish production-scale relevance or latency.

Assessment revision 1 · 2026-09-06 · all five local dimensions passed.

Latest counter-review: 511/511 targeted Rust tests passed for the four modules that were already M4-; the seven newly promoted modules remain bound to their module-specific content-addressed evidence registries.

OPEN LIMITS
  • No production-scale asynchronous ingestion campaign.
  • No representative customer-corpus drift, relevance or latency qualification.
  • No independent assurance opinion.
CURRENT EVIDENCE

What exists today

Structure-aware retrieval, bounded evidence packing and provenance contracts are implemented. The Retrieval crate owns the transient source index exposed as `mentaview_retrieval_source_index` and authenticated `/retrieval/source-index`; it returns metadata and caller-query terms only. The Retrieval page now carries the LongMemEval, LoCoMo, ConvoMem and MemBench comparisons, Qdrant and Milvus assurance, and local index recovery evidence; none is counted as Governed Memory capability. A sealed synthetic 48-query root-to-intent holdout recovered 48/48 expected items with zero wrong-document top-one results and zero scope leaks.

LABEL MEANING

Controlled pilot

A bounded runnable surface exists and can be evaluated in a controlled engagement. Production-scale controls and operating evidence remain gated.

NEXT GATE

What must earn promotion

Qualify production-scale index construction, representative corpus drift, wider prospective holdouts and customer-specific relevance, recovery and latency targets.

Public truth boundary: M4− is a non-standard Mentaview engineering label for internal laboratory validation. It is not an official TRL decision and does not assert production qualification, customer acceptance, independent assurance, certification or universal performance. Inspect the complete assessment record.

← Back to all modulesDiscuss this module