07 / Evidence

Progress measured in holdouts, failure modes and reproducible gates

Review selected Mentaview laboratory evidence, current limitations and the validation discipline used to decide what earns promotion.

All figures below are internal laboratory evidence on stated frozen profiles. They are not independent certification, production SLAs or universal superiority claims.

01

Universal multimodal ingestion controlled pilot

A 735-extension registry, preservation-aware DAG and source-bound checkpoint passed its pinned campaign. Critical audio recovery conditionally reached 55.28% WER and 93.12% coverage on one frozen CHiME-6 conversation. A real fault-injected S01 lifecycle preserved ASR, completed one durable worker attempt and one complete-source diarization request, persisted/read 76 spans and deleted its raw recovery source. Frozen S01 and S21 still rejected the evaluated diarization model at 58.47% and 84.90% DER.

02

Wave 11 question-to-evidence machine review

A sealed 40-case multilingual holdout passed 40/40 through the kernel publication path with zero orphan claims, zero critical false-completes and zero additional model calls; an undisclosed 40-entry blind-review packet was generated.

03

Wave 11 provider-idempotency capability

A sealed qualification passed 12/12 and three loopback HTTP tests passed 3/3, covering explicit header/body placement, retry-key reuse and failed-closed conflicts or unsupported requirements.

04

Wave 11 Rust SDK network contract

Fourteen contract tests passed for an injectable synchronous transport boundary with opaque credential references, bounded request options, explicit readiness and normalized failures.

05

Wave 11 commercial usage reconciliation

Real PostgreSQL and HTTP tests each passed 1/1 for tenant-scoped append-only usage evidence, atomic turn commit, scoped readback and immutable reconciliation snapshots over a closed UTC day.

06

Durable cognitive journey

A 1,000-turn path completed with three restarts on the tested persistence route.

07

External Research production candidate

A frozen 36-case answer-contract holdout passed 36/36 for semantics, provenance and citations, with zero orphan citations, budget failures or research wall-budget exhaustion. Research p95 was 6.891 s and maximum 7.645 s.

08

Embedded semantic retrieval

E5 qint8 reached 28/30 top-1 and 29/30 top-3 versus 11/30 lexical; release p95 was 4 ms on Apple M1.

09

Embedded answer path

A post-hoc corrective replay reached 44/44 assertions with 8 model calls, and reproduced 44/44 on a second model under network denial.

10

Prospective evidence firewall

Three separately executed sealed suites passed 174/174 assertions across the frozen V1, V2 and V3 profiles.

11

Wave 6 transport and media hardening

Connect-time RAG address controls passed 54 adapter tests; bounded multimedia processing passed 96 focused tests, including deterministic rebinding, timeout, excessive-output and kill/reap cases.

12

Wave 7 operational resilience gates

The server admission suite passed 67 tests; an isolated PostgreSQL 17.11 drill restored its table and three-row probe through all five stages; and two network-denied Edge containers preserved exact state while reusing local memory and RAG with zero provider calls observed.

13

Wave 8 shared server admission

A real PostgreSQL test sent 24 concurrent requests through two independent pools and admitted exactly 7 for a limit of 7; tenant isolation, persistence after reconnection, capacity, aggregate usage and pre-authentication state isolation were also checked.

14

Wave 8 point-in-time recovery

A nine-stage PostgreSQL 17.10 physical-backup/WAL drill restored to the inclusive LSN 0/40001C8, retained the before-cut row, excluded the after-cut row, promoted and left zero named Docker resources.

15

Wave 8 sealed offline Edge update

Two content-addressed images crossed an import-only bundle, then baseline → candidate → authorized rollback preserved exact state at turn counts 6 → 10 → 14. Seven tamper, traversal, version, schema and anti-rollback probes failed closed.

16

Reproducible public build

Two isolated production builds matched byte-for-byte across 159 exported files and the paired generated Firebase security configuration.

17

CLI diagnostic-path optimization

The unchanged 344-route local scenario improved from 139.58 s to 117.21 s after successful diagnostics were reused, a 16.0% reduction.

18

Where Mentaview should step aside

On a short-context matrix, direct inference reached 100% in 1.752 s; Mentaview reached 95.71% in 6.113 s.

Promotion protocol

A sophisticated path competes against the simplest valid baseline.

  1. 01

    Freeze the task set before optimization

  2. 02

    Run direct, forced-route and Mentaview paths with comparable models and budgets

  3. 03

    Measure quality, latency, calls, resources, isolation, abstention and failure behavior

  4. 04

    Separate prospective results from post-hoc diagnosis

  5. 05

    Require human review where semantic quality matters

  6. 06

    Promote only the exact profile that passed

Work completed so far

From scaffold to a modular cognitive runtime.

Architecture

Provider-neutral cognitive contracts, separate memory, retrieval, research, ingestion and inference boundaries.

Runtime

Direct Rust path, authenticated tenant-aware server, local/PostgreSQL persistence and controlled provider adapters.

Knowledge

Document ingestion, structural retrieval, evidence graph controls, citations and answer-assurance publication guards.

Continuity

Chat and general memory separation, correction, deletion, restart and isolation journeys.

Edge

Strict composition canary, explicit model packs, semantic retrieval and network-denied replay work.

Assurance

Integrated readiness programs, machine scope, semantic evidence gates, deterministic public builds and action planning.

A deliberate next step

Challenge Mentaview with a baseline that matters to you.

A credible evaluation includes the cases where direct inference, a simpler retriever or no answer at all is the correct result.