The Evaluation and Release Kit activates whenever a module, model, route or deployment profile is proposed for promotion, replacement or customer-facing use.
Evaluation & Release Kit
Make each capability compete against the simplest valid baseline.
Defines frozen holdouts, replay artifacts, acceptance gates, evidence manifests, release checks and capability-specific promotion criteria.
A reproducible decision package that compares the candidate with the simplest valid baseline, records failures and abstentions, binds results to the tested artifact and states whether promotion criteria were met.
Why it exists
The problem this module is designed to solve.
AI features are easy to demonstrate and difficult to qualify. Without frozen baselines, failure accounting and release-bound evidence, teams can promote a complex feature because it looks impressive while ignoring regressions in quality, latency, cost, privacy or abstention.
02 / How it works
A bounded path from need to accountable result.
The public model below describes responsibilities and decisions, not sensitive implementation details, provider secrets or customer data.
- 01
Define the decision
State the capability, baseline, target population, metrics, hard regressions and claim that the evidence is allowed to support.
- 02
Freeze the evaluation
Version holdouts, replay inputs and acceptance thresholds before observing the candidate result.
- 03
Measure the whole trade-off
Evaluate quality together with latency, cost, isolation, failure recovery, abstention and operational constraints.
- 04
Bind evidence to promotion
Record digests, reports and unresolved gates so a passing experiment cannot silently become a broader product claim.
03 / Customer and operator value
Turns quality, latency, cost, isolation, abstention and failure behavior into release decisions instead of marketing intuition.
Baseline-first module evaluation
Frozen prospective holdouts
Failure and abstention accounting
Byte-exact static build checks
04 / Where it creates value
Concrete situations, not generic feature claims.
These are representative product situations. Every deployment still requires its own policy, data boundary and acceptance criteria.
Model or provider upgrade
Prove that a replacement improves the target workload without regressing privacy class, failure behavior or operating cost.
Module release decision
Move a capability from canary to controlled pilot only when its prospective evidence meets the declared gate.
Customer pilot acceptance
Agree on workload-specific success, failure and rollback criteria before using pilot results to justify expansion.
05 / Role in the cognitive system
A clear responsibility creates a trustworthy boundary.
No module is allowed to become an invisible monolith. It owns a narrow contract, composes with named capabilities and refuses responsibilities that belong elsewhere.
What it owns
- Test protocols
- Promotion gates
- Evidence manifests
- Reproducibility checks
What it composes with
What it refuses to own
Boundary before convenience.
Internal evidence is not a certificate, production SLA, legal conclusion or universal superiority claim.
06 / Vision and mission
Evidence over theatre
The kit protects Mentaview's intellectual honesty: capabilities earn their place through reproducible comparison, and the product says exactly what each result can and cannot prove.
AI systems that remain useful, inspectable and sovereign across models, providers and deployment boundaries.
Build the cognitive layer that chooses the smallest sufficient path and turns evidence into accountable action.
Capture the value of advanced AI without surrendering data control, architectural freedom or intellectual honesty.
07 / Evidence and maturity
What the current label means — and what remains open.
Passed with moderate confidence
The evaluation, custody, signature and reproduction tooling is executable and adversarially tested locally; same-process hostility and independent clean-room reproduction remain outside the proven boundary.
- A hostile actor in the same Python interpreter can bypass process-local controls.
- No organizationally independent identities, human gold or licensed-corpus custody campaign.
- No signed clean-room multi-host reproduction or customer acceptance.
What exists today
Frozen holdouts, replay artifacts, acceptance gates and evidence manifests are implemented. Wave 11 adds fail-closed semantic validators for provider idempotency, the injected SDK network contract, the 40-case machine holdout plus blind-review packet, and commercial usage reconciliation. Two isolated public-site builds also matched byte-for-byte across 159 exported files and their generated Firebase security configuration.
Implemented
A software or contract boundary exists in the current development project. This does not by itself mean general availability or independent certification.
What must earn promotion
Add independent review, clean signed-release evidence, production-period operating data and externally agreed customer acceptance protocols before broader assurance claims.
Public truth boundary: M4− is a non-standard Mentaview engineering label for internal laboratory validation. It is not an official TRL decision and does not assert production qualification, customer acceptance, independent assurance, certification or universal performance. Inspect the complete assessment record.