Consolidated qualified baseline
Revision a5efd1b is documented in the public baseline receipt for this sprint.
08 / Roadmap
See what Mentaview is strengthening now, what comes next and how agentic execution and augmented coding remain gated roadmap capabilities.
This roadmap communicates direction, not committed release dates. Roadmap items are not represented as available product capabilities.
Governed documentary assistant
The exact runtime passed 277 targeted local cases. Separately, one actual local browser case completed native TXT upload, a cited local-model answer, placement, reload, placement removal and return to the original stored conversation. This fifth observation preserves four earlier failed observations and is not a reliability estimate.
Results reported . Read the local progress receipt.
Revision a5efd1b is documented in the public baseline receipt for this sprint.
277 targeted runtime cases passed; one separate actual local browser case completed placement, reload, removal and original conversation return; four prior failures preserved. No public backend verification is recorded.
One PostgreSQL case passed with a synthetic provider; default activation remains off. No public backend verification is recorded.
HOLD; hosted backend, representative provider and final release evidence assembly remain open. No public backend verification is recorded.
Release snapshot recorded at 2026-09-14T08:08:42.111099+00:00: canonical publication is M3. Current M8 CI 34815127316 on candidate 93222ec is still running the offline Edge update and rollback gate after application, PostgreSQL/HTTP, tenant-isolation, physical recovery and strict Edge restart passed. Earlier M8 CI 34809614480 remains failed because its manifest expected 23 PITR tests while 33 passed; that exact-count correction is in 93222ec. This new browser-evidence publication is prepared separately and requires its own release qualification. Historical failures remain unchanged. The 277 targeted runtime cases and one successful fifth local browser observation have separate counting boundaries; all four earlier browser failures remain recorded..
Site publication is tracked separately from backend functionality and public activation.
277 distinct cases passed within this targeted run: 44 memory-console Node, 62 pilot-console Node, 169 ordinary server and two separate PostgreSQL cases. The 18 ignored ordinary cases remain outside this execution. Historical 821 runtime tests, the earlier one-case browser observation and independent UI review scenarios are separate evidence, not additional runs of this candidate. The fifth observation supplies one separate actual local browser case; neither its 86 HTTP requests nor the 52 Python/107 Node harness preparation checks are added to the 277 runtime cases.
| Suite | Reported result | Scope and limits |
|---|---|---|
| Memory organization console | 44 passed · 0 ignored Passed locally | Node contract tests cover source selection, explicit wing/room/drawer placement, saved navigation, placement removal, session races and rejection boundaries. |
| Conversation and return navigation | 62 passed · 0 ignored Passed locally | Node contract tests cover pilot-session handling and return to the original conversation. These are UI contract checks, not an actual browser run. |
| Server ordinary library | 169 passed · 18 ignored Passed locally | All-features library tests. The 18 ignored cases retain separate PostgreSQL, attested-sidecar and environment requirements; this run does not clear them all. |
| Guest-memory HTTP with PostgreSQL | 1 passed · 0 ignored Passed locally | One actual HTTP-router test with a fresh disposable PostgreSQL database covers authorized, CSRF-bound, isolated and durable guest memory. |
| Four-layer context HTTP with PostgreSQL | 1 passed · 0 ignored Passed locally | One actual HTTP-router test uses fresh PostgreSQL and a synthetic provider to check four scoped layers, bounded prompt receipts, persisted trace and replay without another generation. Progressive context remains disabled by default. |
Local source boundary. Exact source 274fd7f contains four changed HTML/Node files and one added PostgreSQL integration test: 488 current runtime/Cargo files. Production Rust and Cargo dependency bytes remain unchanged from the earlier grounded composition. All 2,193 archive files were verified before and after qualification. Later site, documentation and CI corrections require separate release checks.
Consent and retention. Explicit consent is required for each public-processing turn, covering the whole conversation and authorized context, including history, summaries, memories, claims and retrieved text. PublicStrict treatment is mandatory; it is not an anonymity guarantee and may reduce detail. Retention is checked for the configured host account: verified none or explicitly acknowledged limited retention. Unknown retention is blocked; no universal provider no-retention claim is made. Creating an offer does not call a provider or decrypt a credential. Authorization is persisted before provider discovery and dispatch; replay does not dispatch again. A new turn needs fresh consent, and expired, revoked or mismatched offers fail closed.
Current runtime qualification. The 277 targeted cases, strict all-features/all-targets server Clippy with the three historical allowances, and the Linux server build passed on the same frozen source. The parent independently verified 32 artifact hashes, the binary readback and absence of both owned databases. This execution contains no actual browser or model call.
Current browser qualification. The fifth separately authorized observation passed one actual local scenario: 86 HTTP requests, one upload, one Ask and zero replay. The pinned local model answered with 17 minutes and R1 in 13,546 ms. PostgreSQL verified initial, saved, removed and returned states, including ordered content digests of both original messages. Removal retained the document; return produced zero mutations. Strict CDP session-completion order, application adoption and independent proxy chronology passed. The original Playwright mixed-origin calculation remains false and diagnostic-only. Parent verified eight original artifacts, three synthetic screenshots and independent database/role absence. Four prior failures remain unchanged; no reliability, representative usefulness or public-backend qualification follows.
Historical qualification. M6 records 821 tests on the earlier runtime; M7 adds one actual local browser case on its earlier binary. Their original receipts remain byte-identical, including pending-at-creation statements. Neither is a re-execution of the current memory experience.
Memory navigation boundary. One actual local synthetic browser case now verifies explicit document placement, a populated drawer after reload, placement removal without document deletion, and return to the same two persisted messages. The 277 targeted runtime cases remain separate. Representative memory usefulness, broader corpora and correction workflows remain unqualified.
Earlier real-model trial. Cedar07 qualified its separate S3 composition with 640 distinct tests and strict Clippy. Cedar03 passed 14 first-turn criteria and 5 replay criteria for one synthetic question with one actual local Qwen generation; zero replay requests were observed at its isolated Ollama HTTP observer. These are separate observations, not a new unique test total. They do not qualify the combined candidate or its browser.
PITR diagnostics. Thirty-three simulated diagnostics tests passed, including a focused negative test for the known generated-secret writer guard. The earlier 32-test proof remains preserved. These mocks do not establish a new physical PostgreSQL PITR success or qualify the integrated release workflow.
Release status. Release snapshot recorded at 2026-09-14T08:08:42.111099+00:00: canonical publication is M3. Current M8 CI 34815127316 on candidate 93222ec is still running the offline Edge update and rollback gate after application, PostgreSQL/HTTP, tenant-isolation, physical recovery and strict Edge restart passed. Earlier M8 CI 34809614480 remains failed because its manifest expected 23 PITR tests while 33 passed; that exact-count correction is in 93222ec. This new browser-evidence publication is prepared separately and requires its own release qualification. Historical failures remain unchanged. The 277 targeted runtime cases and one successful fifth local browser observation have separate counting boundaries; all four earlier browser failures remain recorded.
Historical observations: Pilot foundations verified locally; public activation pending (2026-09-13); Public-processing consent verified locally; activation pending (2026-09-14); Guest-memory UI and PostgreSQL verified locally; public activation pending (2026-09-14); Native TXT/table upload verified locally; public activation pending (2026-09-14); Combined grounded runtime verified locally; browser qualification pending (2026-09-14); One real browser case verified locally; public activation pending (2026-09-14). Earlier results and their source boundaries remain in those dated receipts.
Engineering evidence is not external certification. This prepares evidence for future assurance work. All eleven modules remain at the M4− internal-laboratory boundary, with zero M5 promotions and historical p95 FAIL limits unchanged. No public provider, hosted backend, MVP readiness or S4/S5 closure is claimed. Inspect the existing module assessment.
Truth boundary: Engineering evidence is not external certification. No active milestone is verified live until public qualifying evidence and functional qualification are recorded. Firebase Hosting provenance records static-site publication and does not by itself qualify backend functionality.
NOW
NEXT
NEXT
ROADMAP
ROADMAP
Agentic truth boundary
Generic cognitive planning, budgets, gap-backed escalation, capability boundaries and a developer-assistant contract for workspace identity, code evidence and permissions.
A qualified production tool runtime, autonomous patch execution, language-server integration, terminal coordination, remote repository operations or a benchmarked general agent product.
Prospective task success, regression rate, permission safety, secret exposure, recovery after interruption, useful latency and human acceptance of resulting changes.
A deliberate next step
Roadmap collaboration is welcome when the target workload, authorization model and promotion benchmark are explicit.