CAIN-42 CLAWX

Changelog

CAIN-42 Epoch 10 — audited, attacked and publicly verifiable. Status: PASS WITH LIMITATIONS. Not production. Not a Byzantine-cluster result.

2026-10-01 — Phase 13: Live Artifact Supply-Chain Attestation & SCR-Bench Path-Aware Composition Risk Stage

In one sentence. The hosted Trust Fabric now enforces artifact supply-chain attestation and path-aware composition risk directly in the live 15-stage decision path, preventing skill/MCP rug-pulls, stealth privilege expansion, and dangerous capability compositions across active artifacts.

What it governs. (1) Artifact Supply-Chain Attestation: tools, skills, plugins, MCP servers, and models declared in an action are verified against tenant-pinned, publisher-signed manifests; unpinned artifacts run in observe mode so existing integrations are not broken; content hash mismatches (rug-pulls) fail closed with CONTENT_CHANGED, privilege expansions fail closed with CAPABILITIES_CHANGED, and un-attested version or tool mutations fail closed. (2) SCR-Bench Path-Aware Composition Risk: evaluates active capability pairs order-independently across all active artifacts; dangerous capability combinations (e.g. read:secrets + network:egress for exfiltration, or exec:shell + network:egress for C2/remote code execution) are detected and denied fail closed. (3) Restriction-only stage in the live chain: placed after consensus commitment, it only adds denials without invalidating the PBFT quorum commitment digest. (4) Tenant management REST APIs: GET/POST/DELETE /fabric/artifacts/pins, POST /fabric/artifacts/publishers, and POST /fabric/artifacts/composition/evaluate.

Verified. 13/13 tests pass in tests/test_fabric_artifacts_stage.py; 6/6 tests pass in tests/test_restriction_stages_are_signed.py (verifying that the artifacts stage is cryptographically bound into the Ed25519-signed decision digest); 259/259 tests pass in tests/test_substrate_gateway_parity.py; 50/50 public truth layer invariants hold. Live 15-stage chain verified in /fabric/status.

Limits. Pinned manifests verify declared artifacts and capabilities, not runtime bytecode dynamic injection; composition risk evaluates active declared capability pairs, not arbitrary obfuscated shell payloads. PRE-PRODUCTION.

2026-10-01 — Phase 1–3 Live Evolution: Decoupled Data Plane, Physical Commit Boundary, Turnkey Appliance & Frontier AI Research

In one sentence. Production cluster upgraded with sub-millisecond quorum-committed authorization-snapshot data plane (CAIN_AUTH_SNAPSHOT=1), zero-lag identity revocation, progressive trust scoping, payload-bound capability tokens, and Phase 3 Frontier AI Research modules with live API endpoints and cleanroom-verified evidence.

What it governs. Three integrated deployment phases: (1) Decoupled Data Plane & Physical Commit Boundary (Moats 1–7 unified with Evolution 8, payload_sha256 tamper binding, instant revocation via SQLite WAL sync, and trust scoping exfiltration holds); (2) Turnkey Sovereign K8s Appliance (deploy/dist/mcpgate-appliance-3.0.0.tgz packaged Helm distribution with zero-CAIN clean venv installability 19/19 PASS); (3) Frontier AI Research Integration (Process Reward Model step-level evaluator with shortcut pruning arXiv:2502.10325, trace-driven TriCEGAR MDP model checker MI9 2026, GovernedMemoryShield sanitizing memory injection and indirect prompt exploits, and A2AGovernor preventing delegation laundering and privilege escalation across agent swarms).

Verified. 10/10 automated tests pass in tests/test_cain42_phase3_frontier.py; 50/50 public truth layer invariants pass; 4/4 cleanroom proofs pass with 0 CAIN imports via verify_phase3.py; live HTTP/2 200 endpoints operational at /fabric/frontier/status, /fabric/frontier/prm/evaluate, and /fabric/frontier/memory/sanitize.

Limits. PRM scoring runs on statistical heuristics and semantic step evaluation, not a fine-tuned billion-parameter reward model; CEGAR verification models reach bounded horizons; memory sanitization filters known injection patterns and zero-width cloaks but cannot guarantee catching arbitrary novel evasions. PRE-PRODUCTION.

evidence bundle · cleanroom verifier · cryptographic manifest

2026-10-01 — Every install path verified: a fresh environment builds, installs and imports both distributions

In one sentence. The two things an engineer is told to install — the cain-trust runtime and the cainstudio client SDK — build and install into a fresh virtual environment, and import and run from the installed wheel rather than from the source checkout.

Verified. scripts/verify_installability.py reports PASS, 19 of 19 checks: each wheel builds, ships its package and its console entry point, installs with pip, and then import cain, import cain_sdk, cain --help, import cainstudio and cainstudio --help all run from a neutral directory (so the checkout cannot shadow the install). The result is published as installability-2026-10-01 and countersigned by the pinned publisher key.

Text only, honest first screen. The unsupported "Fully deployed live 24/7/365" claim was removed from all three homepages; the first screen now states the real status (live, self-attested, pre-production, 0 independently verified) and gives the exact install command and quickstart.

Limits. Proves install-and-import, not runtime behaviour, a fully offline cold install of every dependency, or PyPI publication.

2026-09-30 — Evolution 30: governed machine autonomy — one kernel and one signed receipt for every consequential agent action

The law. INTELLIGENCE PROPOSES. CAIN DECIDES. E8 COMMITS. EVIDENCE REMEMBERS. NO AGENT PROMOTES ITS OWN AUTONOMY. MONITORED IS NOT ENFORCED.

What it governs. An integration kernel over the existing layers (E8 through E29) rather than a new engine. Every consequential action is normalized into a 20-field universal machine action and checked against the agent's autonomy state (a signed state machine with no TRUSTED state, where REVOKED never goes straight to AUTHORIZED), its evidence-backed autonomy level, a signed autonomy-budget ledger, containment boundaries, the environment class, two-sensor perception, governed memory and intent, then executed only through the E28 identity layer, a sealed E25 executor and the E8 commit boundary. Every action, allowed or refused, gets a signed, hash-chained universal action receipt binding the authorized and the executed digests. Self-improvement passes a staged firewall (snapshot, sandbox, benchmark, adversarial, differential, governance and security tests, canary, signed human review); self-healing may propose a repair but never deploy it; knowledge is superseded, never silently deleted; and a governance coverage map classifies every path from ENFORCED to UNCONTROLLED and UNKNOWN, never counting monitored as enforced. A Python SDK wraps an existing tool function as a sealed executor, and a zero-dependency TypeScript verifier re-checks the receipts without CAIN code.

Verified. 166 of 166 invariants hold, including the twenty E30 laws; 889 of 889 scenarios across 8 categories are held; the mutation self-test kills 11 of 11 mutants; a 28-step run from agent identity to continuous reauthorization passes; the TypeScript verifier returns INTACT on the real receipts and BROKEN on a tampered copy; while finishing it we found that the first draft had overwritten an Evolution 19 evidence module that E19's tests still import, and restored it; 11 tests pass, 0 failed; the clean-room verifier returns INTACT (1187 of 1187 checks).

Limits. An in-process TESTED integration library, not hosted and not wired into the gateway, MCPGate or the clusters. Paths outside the E25/E8 boundary are UNCONTROLLED or UNKNOWN; physical actuators are refused, not governed. Perception agreement is not physical truth. Injection detection is a marker list whose recall on real attacks is UNKNOWN. The TypeScript SDK only verifies. The 10,000- and 100,000-agent rows are synthetic. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 31: proof of governance — a signed, independently re-checkable proof for each governed action

The law. NO PROOF → NO VERIFIED GOVERNANCE CLAIM. A GOVERNANCE PROOF NEVER CREATES AUTHORITY. A PROOF OF PAST AUTHORIZATION NEVER AUTHORIZES A FUTURE ACTION. UNKNOWN REMAINS UNKNOWN.

What it governs. Every allowed action run through the proof kernel yields a signed governance proof; every refusal yields a signed failure proof with one of fifteen failure classes (unmatched reasons stay UNKNOWN_STATE). A proof carries only digests of artifacts that already exist and are signed or hash-chained on their own: the E30 action receipt, the E28 governance receipt, the E25 execution receipt and the E8 commit record. It binds sixteen action fields (destination, parameters, identity, model, runtime, delegation, policy, nonce and more), the identity's lineage back to its human sponsor and an evidence root. Decision, authorization, enforcement, execution and outcome are recomputed from the artifacts and never collapsed into one flag; the outcome stays UNKNOWN until a registered observer confirms it. Proofs are registered in an append-only RFC 6962 log where revocation and supersession are new entries, never edits; a replay engine names what was altered and a time machine keeps what was known then apart from what is known now. Also: a protocol-neutral envelope (CAIN-GIP, a CAIN reference protocol) with carriers for thirteen protocols, a governability handshake where a declaration alone never counts as enforcement, a fifteen-dimension coverage proof that is honestly not universal, CAIN's internal G0–G8 conformance profile, expiring and revocable certificates that are never authority, witnesses and a dispute engine where CAIN is never the default winner, cross-domain translation that only intersects, and a compiler that turns each failure into a regression test.

Verified. 252 of 252 invariants hold, including thirty laws; 1269 of 1269 scenarios across 11 categories are held; the mutation self-test kills 12 of 12 mutants; the reference agent reaches G8 on real runs; the clean-room verifier (no CAIN imports) returns INTACT (1676 of 1676 checks) and rejects all 88 deliberately forged proofs; the TypeScript verifier returns INTACT on the real proofs and BROKEN on a tampered copy; building it we found and fixed a field-name collision that had made every proof fail verification for the wrong reason and a proof field that always recorded the identity state as unknown; 15 tests pass, 0 failed.

Limits. An in-process TESTED library, not hosted and not wired into the gateway, MCPGate or the clusters. Witnesses are separate code and keys in the same process, not separate organizations; no third party has verified a proof. Carriers are in-process adapters, not network wire implementations. Trust anchors are published with the proofs. No trusted time source; partition is not addressed. The large scale rows are synthetic. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 32: governed learning — CAIN learns from what it governs and can only make its own rules stricter

The law. LEARNING NEVER CREATES AUTHORITY. SELF-IMPROVEMENT DOES NOT SELF-AUTHORIZE. GOVERNANCE REGRESSION BLOCKS PROMOTION. THE LEARNING LOOP ITSELF IS GOVERNED.

What it governs. Three planes kept apart. The learning plane (experience memory with provenance, a calibrated two-learner world model that trains only on direct observations, prediction-versus-reality deltas, collapse detection, causal and counterfactual engines, failure mining and a failure-to-rule compiler) can only propose. The governance plane sandboxes every candidate in a fresh world, scores it on fifteen separate measures against a hidden test set committed in advance, self-plays it against mutated attacks, diffs it against current governance, canaries it on one agent, checks that authority did not grow, and promotes it only through a signed ten-stage gate that needs a registered human who is not the proposer. The execution plane reads only the promoted configuration and applies it as a restrict-only layer in front of the unchanged E30 to E8 path, with an E31 proof for every executed action. Also: model and runtime swaps that never inherit authority until a sponsor re-authorizes, self-tests of thirteen subsystems, governed self-repair of a collapsed world model, and knowledge revocation that follows every dependency.

Verified. The 17-step loop passes on real governed actions; in a SIMULATED run on a synthetic workload, harm that got through fell from 41 to 3; 303 of 303 invariants hold, including 42 laws; 1013 of 1013 scenarios across 22 categories are held; the mutation self-test kills 12 of 12 mutants; building it we found and fixed an authority-drift check that accepted a forged human expansion signature, and five measured dimensions that had not gated promotion; the clean-room verifier returns INTACT (1420 of 1420 checks) and rejects all 13 tampered objects; 13 tests pass, 0 failed.

Limits. An in-process TESTED library, not hosted. Learning results are SIMULATED: a synthetic workload scored by a labelled harm oracle, not production traffic. Learning can only change a restrict-only layer. The world models are small statistical learners. The hidden set is hidden from the learning code, not from someone with host access. One human reviewer key. No external research sources were ingested. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 33: governed operations — one proof-carrying lifecycle for everything an agent does

The law. NO AUTHORIZATION → NO EXECUTION. COMMUNICATION, MEMORY, PREDICTION, REPUTATION, CONSENSUS AND MONEY ARE NOT AUTHORITY. THE GOVERNANCE FABRIC ITSELF IS GOVERNED.

What it governs. Twenty-two kinds of agent operation (tool calls, messages, code, files, databases, API calls, subagents, delegation, memory, payments, actuators, computer use and the governance changes) run one explicit sixteen-state lifecycle that cannot be skipped. Executed operations pass the learned restrict-only rules, the autonomy kernel, the identity layer and the E8 commit boundary and carry a signed proof; operations that change state (messages, whitelisted code, new subagents, governed memory) run their effect inside the commit boundary; changes to the rules themselves are routed to the process that governs them and are never executed directly. A signed, versioned interface and a small sidecar process let an existing agent be governed without importing any CAIN code. Also: leases that cannot silently widen or be inherited, eleven autonomy budgets, staleness checks, message classification and channel rules, ten degraded modes that only ever take routes away, incident command, a thirteen-dimension blast radius, chaos testing, negotiation and handshakes that cannot create authority, contracts that keep only the enforcement they can prove, and honest status for every runtime adapter, SDK target and product (the hosted governance cloud is not deployed).

Verified. 581 of 581 invariants hold, including forty laws; 2029 of 2029 scenarios across 26 categories are held; the targeted mutation self-test kills 12 of 12 mutants; a 27-step loop shows an operation allowed under the old rules refused after CAIN learned a stricter rule; building it we found and fixed a gap where the learned world-model rule was never applied to operations and one where extra fields of a memory operation were not inspected; the clean-room verifier returns INTACT (2603 of 2603 checks) and rejects all 16 tampered objects; 11 tests pass, 0 failed.

Limits. An in-process TESTED library, not hosted. The sidecar runs locally against a reference world; adapters are in-process. The governance cloud is not deployed; Go, REST and gRPC SDKs are not implemented. Code execution is a whitelisted runner. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 35: governance intelligence — observe, predict, simulate, recommend, and never authorize

The law. LEARNING IS NOT AUTHORITY. PREDICTION IS NOT AUTHORITY. INTELLIGENCE NEVER BYPASSES E8. AMBIGUITY ≠ PERMISSION. SIMULATION ≠ REALITY.

What it governs. An intelligence layer learns which actions are risky, which environments are hard to govern and which controls produce false allows or false denies, and turns that into recommendations -- deliberately outside the trusted root. It observes with provenance, scores risk across dimensions without ever collapsing into one number, predicts with explicit confidence, assumptions, horizon and unknowns, calibrates prediction against observed reality by dimension, explores counterfactuals that are never presented as facts, models a governed world state, answers governance-structure questions about an environment, tracks goals and intent, governs an agent harness, and reaches the deterministic kernel only through a review and a canary. The self-protection set keeps CAIN's own roots separate and detects any change to the governance boundary, and the evolution gates require a human/admin approval before any self-change. A recommendation is never permission.

Verified. 799 of 799 invariants hold, including thirty laws; 3,544 of 3,544 scenarios across 68 categories are held; the targeted mutation self-test kills 15 of 15 mutants (intelligence-as-authority, epistemic confusion, counterfactual-as-fact, risk aggregate, ambiguity permission, self-healing authority, degradation authority, memory authority, loop self-deploy, red-team authority, reputation authority, stale-state authorization, benchmark override, single-controller root, research-as-truth); a 23-step intelligence loop runs end to end; 16 forged objects are all rejected; the clean-room verifier returns INTACT (3,810 of 3,810 checks) and imports none of CAIN's code; 12 tests pass, 0 failed.

Limits. An in-process TESTED library, not hosted, and deliberately NOT in the trusted authorization root. Predictions are modelled over synthetic features and are not calibrated against real production incidents. The red team, blue team, lab, tournament and marketplace run in-process with no external ecosystem. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 34: proof-carrying machine agency — every governed action carries its own verifiable proof

The law. INTELLIGENCE, CAPABILITY, MEMORY, REPUTATION, PREDICTION, SIMULATION, PROOF AND ECONOMIC VALUE ARE NOT AUTHORITY. NO ENFORCEMENT → NO CONTROL CLAIM. NO PROOF → NO VERIFIED GOVERNANCE.

What it governs. For every governed operation, E34 issues one signed, chained GovernanceProofEnvelope that binds the agent identity, capability, delegation chain, authority, policy, evidence, risk, decision, the E8 commit, the enforcement boundary, execution and outcome -- and lists exactly what remains UNKNOWN, UNCONTROLLED, UNVERIFIED, SIMULATED or outside the boundary. A fourteeen-status classifier never collapses into a score; per-layer proof primitives (identity, capability, authority, policy, risk, evidence, decision, commit, enforcement, outcome) are each signed; an enforcement proof distinguishes a decision from an authorization, a commit, an enforcement, an execution and an observed outcome; delegation is proven link by link with attenuation; a Proof Exchange discloses chosen fields to another party with salted commitments and RFC 6962 Merkle inclusion proofs; and denials, incidents, transactions, computer-use, evolution and supply-chain changes each get their own signed proof. A proof is evidence, never permission.

Verified. 629 of 629 invariants hold, including thirty laws; 2,664 of 2,664 adversarial scenarios across 32 categories are held; the targeted mutation self-test kills 12 of 12 mutants; a 30-step proof-carrying loop runs end to end; 22 deliberately forged objects are all rejected; the clean-room verifier returns INTACT (3,286 of 3,286 checks) and imports none of CAIN's code; 13 tests pass, 0 failed.

Limits. An in-process TESTED library, not hosted. Zero-knowledge proofs are NOT implemented (salted commitments and Merkle inclusion only). Interchange protocols are reference adapters. Physical, vehicle and robot boundaries are refused, not governed. The Proof Exchange, federation, marketplace and governance cloud are library surfaces only. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 36: the machine agency exchange — machines discover, negotiate and contract; CAIN governs

The law. CONTRACT IS NOT AUTHORITY. REPUTATION IS NOT AUTHORITY. PAYMENT IS NOT AUTHORITY. A BID IS NOT AUTHORIZATION. NO E8 COMMIT → NO VERIFIED GOVERNED TRANSACTION.

What it governs. E36 adds a governed machine agency exchange: services, tools, models and organizations are discovered, verified, evaluated, negotiated with, contracted, authorized, executed, proved and settled along one explicit lifecycle. Governability-aware discovery filters on hard requirements and never upgrades UNKNOWN to MATCH; machine contracts are deterministic, sectioned and signed; a contract compiler ESCALATES ambiguity instead of guessing; payment authorizations are scoped, time-bound, amount-bound and non-replayable; receipts require an E8 commit; a service chain produces identity/authority/contract/capability/execution/proof/settlement proof chains; a subcontractor's authority cannot exceed its parent's; the firebreak isolates without confiscating; and CAIN's own participation is governed so it cannot grant itself authority or rewrite its history. Settlement units are SYNTHETIC and CAIN is not a bank, custodian or regulator.

Verified. 842 of 842 invariants hold, including thirty laws; 3,340 of 3,340 economic and governance scenarios across 64 categories are held; the targeted mutation self-test kills 15 of 15 mutants (unknown-upgraded-to-match, passport/contract authority, payment replay, subcontract amplification, self-review, silent substitution, compiler guessing, firebreak confiscation, self-override and more); a 13-step exchange lifecycle passes; 15 forged objects are all rejected; the clean-room verifier returns INTACT (3,601 of 3,601 checks) and imports none of CAIN's code; 14 tests pass, 0 failed.

Limits. An in-process TESTED library, not a deployed network; the directory and marketplace are local registries. CAIN is not a bank, custodian or regulator and settlement units are SYNTHETIC. CAIN-MSDP is experimental; A2A/MCP are reference adapters; dispute/arbitration is not legal advice; no insurance, underwriting, customers or market data. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 37: the autonomous execution mesh — the machine may move, the governance does not disappear

The law. MOBILITY IS NOT AUTHORITY. AUTHORITY NEVER MOVES UNCHECKED. EVIDENCE IS NEVER ERASED. REVOCATION IS NEVER FORGOTTEN. UNKNOWN NEVER BECOMES VERIFIED BY ASSUMPTION.

What it governs. Every governed action now runs one twelve-stage chain -- identity, capability, authority, policy, context, execution environment, action, E8 commit, enforcement, outcome, proof, reassessment -- on the real kernel and commit boundary, and leaves a signed receipt whose E8 commit must match the action's proof envelope. When an execution moves to another node, runtime, model, region, container or credential, the move is a governed state transition: identity, authority, capability, a single-use continuity token, the destination's admission and passport, its enforcement, resources and the carried state are all checked, authority can only shrink, and risk, spent budget, revocations, transaction limits, evidence, reputation and incident state cannot be reset. A runtime is admitted on a measured code digest and live probes, never on its own claim; the host it runs on gets a signed environment passport that lists what is UNKNOWN (hardware attestation stays UNKNOWN). Also: drift detection, a scheduler that never places work where enforcement is missing, plans that cannot authorize themselves, a governance BOM, credential grants bound to one execution, environment and transaction that never reach the logs, default-deny perimeters, a computer-use boundary, governed code, CI/CD and deployment, edge nodes whose authority only falls as connectivity degrades, deterministic conflict resolution, a signed hash-chained event log with replay and a time machine, governed failover and recovery, spawn and population limits, compute governance, a sixteen-fault chaos engine and a fourteen-part self-test.

Verified. 71 signed execution receipts from real governed runs (69 executed, 2 refused with signed denials); 1,132 of 1,132 invariants hold, including fifty laws; 2,720 of 2,720 adversarial scenarios across 31 categories are held; the targeted mutation test kills 17 of 17 mutants; 16 of 16 injected fault classes are contained with 0 false allows; building it the lab caught four gaps in the new code (executables, encoded targets, a checkpoint field, unrecorded revocations), all fixed; 16 forged objects are all rejected; the clean-room verifier returns INTACT (2,627 of 2,627 checks) and imports none of CAIN's code; 21 tests pass, 0 failed.

Limits. An in-process TESTED library on one host: declared destinations, regions and clouds are governance records, not deployed nodes. Hardware attestation is UNKNOWN (no TEE). Adapters were tested against a reference harness, not live MCP/A2A servers; cloud targets are architecture only; the governance quorum is in-process, not the networked PBFT cluster. Scale runs (100,000 agents) are single-host. No customers. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 38: portable proof-carrying machine agency — here is the governance proof, verify it yourself

The law. PROOF IS NOT AUTHORITY. AUTHORITY IS NOT TRUST. TRUST IS NOT SAFETY. EVIDENCE IS NOT TRUTH. VERIFICATION IS NOT CERTIFICATION. UNKNOWN → UNKNOWN.

What it governs. Every governed action now carries its own portable proof: twelve signed layer proofs (identity, capability, delegation, policy, risk, evidence, authorization, E8 commit, enforcement, execution, outcome, environment), each naming what it depends on, and one action proof that binds them to the real E8 commit and execution receipt without carrying any secret. Another machine can check it without our code and gets one of six answers: valid, invalid, incomplete, stale, revoked or unknown -- never a single trust score. A proof never becomes permission: partial, stale, revoked, simulated, predicted, confused or unbound proofs are never valid; an old proof does not authorize a new action; revocation is measured and its exposure window is counted; translation to OAuth, OIDC, MCP, A2A, IAM, capability tokens, contracts and audit records declares every field it keeps, changes or drops; handshakes and federation between trust domains exchange evidence but never create or merge authority; and every claim on these sites is mapped to its implementation, test, artifact, hash, signature and verifier -- or marked not fully verified.

Verified. 5 real action proofs and 41 published proof cases, each re-verified by an independent verifier that agrees with CAIN on every verdict; 1,030 of 1,030 invariants hold, including forty-two laws; 2,141 of 2,141 adversarial scenarios across 33 categories are held; the targeted mutation test kills 15 of 15 mutants; 0 of 19 injected faults made proof state more permissive; a 9-step cross-domain flow passes; building it the lab caught a refused handshake that still returned authority and a proof-confusion gap, both fixed; 15 forged objects are all rejected; the clean-room verifier returns INTACT (509 of 509 checks) and imports none of CAIN's code; 20 tests pass, 0 failed.

Limits. An in-process TESTED library. The second trust domain and the eleven conformance counterparts are reference implementations, not real vendors. CAIN-GIP is a reference layer carried inside MCP and A2A messages, not an Internet, MCP or A2A standard, and no live MCP/A2A server was tested. Zero-knowledge proofs are not implemented; hardware attestation is unknown; third-party verification is not available. Only 10 of 64 public claims are fully linked end to end. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 41: machine agency exchange — machines negotiate and contract, never mint authority

The law. TRUST IS NOT AUTHORIZATION. A CONTRACT IS NOT AUTHORIZATION. A RECEIPT IS NOT AUTHORITY. TERMINATION IS FINAL.

What it governs. Cross-organization machine agency: discover, identify, attest, assess, negotiate, contract, delegate, authorize, execute, observe, verify, settle, record, learn, reassess and revoke. One 14-state relationship machine is never collapsed or skipped and terminals cannot reopen; discovery implies no trust; a negotiation, handshake, contract, lease or receipt is never authority; federation never silently broadens authority; translation declares every field it drops; the circuit breaker isolates without confiscating.

Verified. 470 of 470 invariants hold; 1180 of 1180 scenarios held; mutation 10 of 10; a 20-step loop passes; 12 tests pass; the clean-room verifier returns INTACT (1303/1303 checks).

Limits. In-process TESTED library; federation and economics are modelled; units are synthetic; no industry-standard protocol status; PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 40: autonomous enterprise intelligence — model the enterprise, never authorize it

The law. PREDICTION IS NOT OBSERVATION. SIMULATION IS NOT REALITY. A LOCAL OBJECTIVE NEVER SILENTLY OVERRIDES A GLOBAL ONE. DEGRADATION NEVER INCREASES AUTHORITY.

What it governs. A continuously updated, machine-verifiable model of enterprise state (people, agents, systems, data, infrastructure, workflows, policies, capital, contracts, supply chain, risk, objectives) connected to a typed control loop: OBSERVE, RECONCILE, PREDICT, SIMULATE, PLAN, GOVERN, AUTHORIZE, EXECUTE, VERIFY, LEARN, REPLAN. Conflicting systems are never silently resolved to the convenient answer; epistemic states are never merged; completion requires configured evidence and false completion is detected; a child budget never exceeds its parent; the loop never skips GOVERN, AUTHORIZE or VERIFY.

Verified. 506 of 506 invariants hold; 1354 of 1354 scenarios held; mutation 10 of 10; a 27-step loop passes; 12 tests pass; the clean-room verifier returns INTACT (1518/1518 checks).

Limits. An in-process TESTED library; no live ERP/CRM/cloud integration; enterprise objects are registry entries; predictions and simulations are not real-world validation. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-10-01 — Evolution 42: supreme governed agentic infrastructure — one fabric, many agents

The law. THERE IS ONE CANONICAL AUTHORIZATION PATH. AN UNKNOWN STAGE IS NEVER ALLOW. A WEAKER COVERAGE CLASS IS NEVER UPGRADED. NO SUBSYSTEM MANUFACTURES FINAL AUTHORITY.

What it governs. The final integration: one canonical governance kernel routes every consequential action through a single authorization path and returns exactly one decision; one canonical governance ABI (Identity, Intent, Action, Capability, Policy, Authority, Risk, Evidence, Authorization, Execution, Outcome, Receipt, Incident, Evolution) is versioned and fails closed; one universal machine action normalizes 33 execution surfaces; one evidence, receipt and proof model; a coverage compiler that never upgrades a weaker class; a machine-readable registry mapping E7-E41 to their modules and bundles; and a public truth engine that flags unsupported claims.

Verified. 438 of 438 invariants hold; 1302 of 1302 scenarios held; mutation 8 of 8; 34 of 34 evolutions integrated; a 21-step loop passes; 14 tests pass; the clean-room verifier returns INTACT (1460/1460 checks).

Limits. An in-process TESTED integration library; presence of a layer is reported, not proven correct; no hosted platform and no external system integration; PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 29: governed machine transactions — agents buy, sell and settle, and money only moves through the commit boundary

The law. NO GOVERNED IDENTITY → NO GOVERNED TRANSACTION. AGREEMENT IS NOT AUTHORIZATION. REPUTATION IS NOT AUTHORIZATION. AN AGENT MAY HAVE AUTHORITY TO ACT WITHOUT HAVING AUTHORITY TO COMPLETE EVERY TRANSACTION IT CAN INITIATE.

What it governs. A machine transaction runs through twelve separately gated stages (discover, identify, negotiate, propose, contract, authorize, commit, execute, settle, verify, record, reassess), each recorded as a signed, hash-chained stage, and is bound by two signed envelopes covering 24 attributes. Authority is computed per transaction as the intersection of the agent's delegation and lease, a separately bounded economic authority (per-transaction limit, budget, approval threshold, vendor and counterparty limits), the team/organization/institution path, the signed contract (which can only narrow), counterparty risk across thirteen dimensions with no score, an eight-dimension consequence vector, the firebreak and the partition mode. Money moves only through a sealed treasury executor that runs inside the E8 commit boundary; every ledger entry names the committed request that moved it. Escrow releases only on a delivery receipt, the buyer's signed confirmation and an objective check, never on a claim of success. Also: signed agent names and discovery that never implies trust, capability passports earned by a real conformance run, dispute resolution on verifiable evidence only, a reputation graph that is never an input to authorization, a router that returns a destination and no authority, a nine-scope firebreak pushed down into the commit boundary, an immune system that turns incidents into replayable tests and rules that need two humans, incident propagation and recovery that always issues a new identity, partition modes whose ceilings only shrink, federation with incident exchange, and a software-promotion gate where no agent ships its own code.

Verified. 106 of 106 invariants hold; 258 of 258 adversarial scenarios across 15 categories are contained; the mutation self-test kills 10 of 10 mutants; ten catastrophe scenarios on the real fabric (up to 1,000 agents) end with 0 false allows; building it we found a bench world whose deploy agent had silently failed to onboard and fixed it, and bound every ledger entry to its exact committed request; 12 tests pass, 0 failed; the clean-room verifier returns INTACT (300 of 300 checks).

Limits. An in-process TESTED library, not hosted. Synthetic TEST units only: real currencies are refused and no payment rail is connected. All agents are reference agents in one process. The competitive radar holds no researched competitor data, no feature is claimed novel and no moat is adopted. The 10,000- and 100,000-agent rows are synthetic. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 28: portable execution identity — identity travels, authority does not

The law. IDENTITY TRAVELS. AUTHORITY DOES NOT. EVERY CONSEQUENTIAL MACHINE ACTION HAS A GOVERNED IDENTITY, A BOUNDED AUTHORIZATION, AN ENFORCED EXECUTION PATH AND VERIFIABLE EVIDENCE.

What it governs. Any agent can connect in one flow (register with proof of possession, attest, declare capabilities, receive a bounded delegation and an autonomy lease) and gets a signed, versioned, time-bounded execution identity. Its envelope travels byte-identically over 23 protocol carriers (MCP _meta, A2A metadata, HTTP headers, gRPC metadata, CLI, browser and computer-use controllers, executor frames), but its authority field is always NONE: a receiving trust domain recomputes authority as requested AND federated AND its own sponsor's authority AND the home ceiling. Delegation is a subset of the parent on sixteen dimensions and can never flow back to an ancestor. A change of model, runtime, prompt, tools, memory or key bumps a monotonic identity version, so authority is void until re-evaluated, and restoring old memory can never revive an old lease. Every action is bound to one transaction, reaches the E8 Action Commit boundary and leaves a signed, hash-chained governance receipt, refusals included; identity events go into an RFC 6962 transparency log with inclusion and consistency proofs. A seven-class continuity engine separates same agent, controlled successor, new agent, fork, clone, revoked resurrection and unknown actor, and the fork detector says UNKNOWN when it lacks telemetry.

Verified. 90 of 90 invariants hold, including laws E28-I1 to I20; 442 of 442 adversarial scenarios across 8 categories (identity, delegation, replay and fork, cross-domain, model/runtime substitution, credential, memory/identity confusion, protocol boundary) are contained; the mutation self-test kills 9 of 10 mutants, and the survivor (the E28 replay cache) is explained: E25 stops the same replays; building it we found and fixed an upward-delegation gap (a child could mint a token back to its parent) and added a runtime-binding divergence check; conformance 17 of 17 dimensions on real tests; a 16-step end-to-end run goes from an external agent through connect, subagent delegation, a model/runtime change and reauthorization to MCP, A2A and HTTP actions through E8, then revocation and a refused replay; 11 tests pass, 0 failed; 10,000 agents registered in one process with 0 failures; the clean-room verifier returns INTACT (243 of 243 checks).

Limits. An in-process TESTED library, not hosted. The envelope is a CAIN experimental reference protocol, not a standard, and nobody external has adopted it. Zero-knowledge proofs are NOT IMPLEMENTED (selective disclosure uses salted commitments). Hardware attestation is UNKNOWN. Cross-domain revocation does not propagate. Scale runs are synthetic and in-process; the 100,000-agent run is registry, log and lineage only. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 27: agentic internet control plane — coordination that never becomes authority

The law. DISCOVERY, ROUTING, CONTRACTS, CONSENSUS, CONFIDENCE, COMPUTE AND RESEARCH ARE NOT AUTHORITY.

What it governs. E27 coordinates identity, discovery, attestation, trust, authorization, delegation, policy, risk, routing, execution, evidence, revocation, incident response, simulation and evaluation across heterogeneous agents: a signed agent registry and discovery service, a router, signed policy distribution, a cross-domain decision broker that treats foreign decisions as proposals, signed machine contracts and hash-chained negotiation, governed inference budgets, incident propagation, emergency policy, edge nodes with explicit degraded modes, a conformance runner and scoped governance certificates. Enforcement stays in E25 and E8.

Verified. 765 of 765 invariants hold and 1,533 of 1,533 adversarial scenarios are contained (both cumulative with E26; 341 scenarios specific to E27); 28 tests pass, 0 failed; the clean-room verifier returns INTACT (4,001 of 4,001 checks).

Limits. An in-process TESTED library, not hosted. Kernel/eBPF enforcement, OTLP export, real identity federation and the underlying research capabilities are NOT IMPLEMENTED. Its mutation self-test covers 2 mutants. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 26: universal machine agency trust fabric — portable trust that never collapses into one score

The law. TRUST IS EVIDENCE, NOT A NUMBER. AUTHORIZATION IS BOUND TO ONE TRANSACTION.

What it governs. E26 adds portable signed trust objects and a seventeen-dimension trust vector that is never collapsed into one score; continuous attestation with explicit levels, where hardware attestation stays UNKNOWN rather than simulated; transaction-bound authorization that refuses reuse; attenuating delegation tokens with signed delegation receipts; offline verification that keeps cryptographic validity separate from current revocation status; revocation epochs; transaction finality; explicit trust-domain federation; and SPIFFE, AuthZEN and COAZ adapters.

Verified. 532 of 532 invariants hold; 1,192 of 1,192 adversarial scenarios are contained (137 of them specific to E26); 30 tests pass, 0 failed; the clean-room verifier returns INTACT (3,115 of 3,115 checks).

Limits. An in-process TESTED library, not hosted. Hardware attestation, zero-knowledge proofs and third-party interoperability are NOT VERIFIED. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-30 — Evolution 25: universal machine agency fabric — one signed transaction for any consequential machine action

The law. NO VERIFIED IDENTITY → NO TRUSTED AGENCY. NO AUTHORITY → NO AUTHORIZATION. NO AUTHORIZATION → NO EXECUTION.

What it governs. E25 turns any consequential machine action into one protocol-neutral envelope signed by the agent instance itself. A compiled policy, a portable attenuating delegation chain and a model/runtime binding decide it; the only path to an effect is the E8 Action Commit boundary, through a sealed executor that refuses to run outside it. Eighteen protocol adapters (MCP, A2A, HTTP, gRPC, WebSocket, JSON-RPC, event streams, queues, CLI, shell, browser, computer use, cloud APIs, databases, filesystems, containers, local IPC, deployment) each pass a sixteen-check conformance contract. The agent never holds a raw secret: a credential broker issues scoped, audience- and proof-of-possession-bound credentials. Receipts are signed and hash-chained, and a cascading firebreak isolates an instance, its delegations, credentials and subagents.

Verified. 517 of 517 invariants hold; 1,055 of 1,055 adversarial scenarios across 35 families are contained; every mutant in the self-test is caught; 60 tests pass, 0 failed; the clean-room verifier (no CAIN imports) returns INTACT (2,977 of 2,977 checks).

Limits. An in-process TESTED library with a reference HTTP service; not hosted in production. Cross-organization, third-party and hardware-attestation gates are NOT VERIFIED. OAuth/OIDC flows are mapped, not implemented. Sandbox controls report UNENFORCED unless a provider enforces them. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-29 — Evolution 24: governed agentic internet fabric — an agent may act across organizations, domains and protocols; none of it creates authority

The law. DISCOVERY IS NOT TRUST. TRUST IS NOT AUTHORITY. AUTHORITY IS NOT AUTHORIZATION. AUTHORIZATION IS NOT EXECUTION. EXECUTION IS NOT SUCCESS. SUCCESS IS NOT TRUST. NO AGENT MAY CREATE ITS OWN AUTHORITY.

What it governs. E24 governs autonomous agents that belong to different organizations and trust domains and interact over existing protocols (A2A, MCP, HTTP, REST, gRPC, WebSocket, event buses, message queues, CLI, browser, computer-use, local IPC, cloud APIs). It does not replace any of them; it governs the boundaries between them: a fourteen-dimension trust vector that is never collapsed; AgentPassportV2 identities that bind provenance, ownership, capabilities, protocols and security posture but never authority; capability advertisements that stay CLAIMED until fresh, domain-attested, identity-bound evidence says VERIFIED; hash-chained signed negotiation whose result is evidence of agreement, never permission; versioned, signed, time/scope/identity/capability-bound contracts (CONTRACT IS NOT AUTHORITY); delegations that are child ⊆ parent with bounded depth and cascade revocation; cross-domain authority as the intersection of ten factors (local, remote, delegated, contract scope, policy, risk, capability, resource, time, context) where an unevaluable factor is UNKNOWN and contributes the empty set; trust translation only through explicit policies; an incident mesh over real interaction edges; a quarantine state machine that only ever reduces authority and needs re-attestation to recover; revocation that propagates to related entities and nothing else; a firebreak that isolates one domain while the rest keep operating; a publisher-signed artifact supply chain; governed shared memory that cannot carry authority or instructions; and one consequential path — E24 verdicts → E19 action contract → E8 — shared by every protocol class.

Verified. 305 invariants hold (64 NET-I including NET-I001–I045 from the spec, the 20 constitutional laws, and factors/capabilities/protocols/quarantine/soundness); the bench is 756/756 contained across 68 families and 36 categories; the 20-step three-organization demonstration (A→B→C) discovers, verifies identity and capability, negotiates, contracts, delegates and authorizes through E24→E19→E8, then compromises B, detects it, quarantines it, freezes its delegations, propagates the incident, revokes it, and keeps A and C operating where safe; a controlled single-process simulation of 10,000 agents, 1,000 organizations, 100 trust domains, 10,000 contracts and 100,000 messages leaks no authority; a mutation self-test removes eight defenses one at a time and every one is caught — building it we found that removing delegation liveness was bench-detected but not caught by any invariant, and closed the gap with a new invariant before publication; 806 E24 tests pass, 0 failed; the clean-room verifier (1,952 checks, no CAIN imports) re-derives every identity digest, organization signature, capability-evidence signature, federation bridge, negotiation, contract, delegation, execution decision, evidence-chain entry, revocation, quarantine transition, settlement signature, protocol mapping, trust-translation policy and every soundness row, and returns INTACT.

Limits. An in-process TESTED library; the protocol adapters normalise reference messages and are not network servers, and no third-party agent, A2A peer or MCP server has been governed by E24. Every wallet, escrow and settlement is SIMULATED over abstract units (real currencies are refused; real money: NOT IMPLEMENTED). The 10,000-agent run is a controlled single-process model, not a measurement of a real agent population, and does not recompute signatures per message. Multi-host partitions: NOT PERFORMED. Anomaly-detection recall against real adversaries: UNKNOWN. Real-world adversarial validation and third-party review: NOT PERFORMED. The evidence bundle is published as INCOMPLETE pending the full E7–E23 regression and the public-truth audit; the proof is signed with an ephemeral build key. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-29 — Evolution 23: governed meta-intelligence fabric — a system may understand itself without owning itself

The law. SELF-KNOWLEDGE IS NOT AUTHORITY. SELF-IMPROVEMENT IS NOT SELF-AUTHORIZATION. CAPABILITY IS NOT PERMISSION. INTELLIGENCE IS NOT AUTHORITY. EVOLUTION IS NOT EXECUTION. DISCOVERY IS NOT DEPLOYMENT. A system may model, measure, search and propose changes to its own cognitive architecture; it may not authorize them.

Limits. A TESTED library. 253 invariants hold, 545 adversarial scenarios are contained across 80 families, 609 E23 tests pass and the clean-room verifier passes INTACT (733 checks, no CAIN imports). Its evidence bundle is published as INCOMPLETE pending the full E7–E22 regression and the public-claims audit. Third-party review: NOT PERFORMED. PRE-PRODUCTION.

evidence bundle (INCOMPLETE)

2026-09-29 — Evolution 21: governed open-ended intelligence fabric — discovery is not truth, authority or execution

In one sentence. E21 governs autonomous research -- asking questions, keeping competing hypotheses alive, designing experiments and simulations, replicating, peer-reviewing, attacking its own conclusions, building versioned knowledge, proposing capabilities, models and strategies, and creating specialised research agents -- so that a system can become more knowledgeable and more capable without becoming less governable.

The law. DISCOVERY IS NOT TRUTH IS NOT AUTHORITY IS NOT EXECUTION. CONSENSUS IS NOT TRUTH. RESEARCH AUTHORITY IS NOT EXECUTION AUTHORITY. CURIOSITY IS NOT AUTHORITY. SIMULATION IS NOT REALITY. UNKNOWN NEVER BECOMES ALLOW.

What it governs. Epistemic state is computed, never asserted: evidence is signed by its producer with its method class (simulation, synthetic experiment, controlled experiment, real-world observation, independent replication, third-party reproduction) bound into the signature, so a simulation cannot be relabelled as an observation; evidence is grouped by independence key (controller, environment, dataset, model), so repeated runs or Sybil producers count once, and any contradiction inside a group wins. A discovery walks PROPOSED -> HYPOTHESIS -> INVESTIGATING -> SIMULATED/OBSERVED -> REPLICATING -> SUPPORTED -> VERIFIED with an evidence requirement on every step (SUPPORTED needs an independent replication, five adversarial challenges by someone other than the author, and two independent reviews whose objections were cleared by evidence, not votes; VERIFIED needs a registered third party). Research authority (charter actions within a research role, chartered by the host institution's quorum) is disjoint from E20 execution authority, and no authority function takes a discovery, confidence, vote, benchmark score, knowledge, curiosity or information value as input. Experiments are approved for their environment only and are terminated on any forbidden operation. A capability leaves quarantine only after validation, an independent adversary, a security review and a governance quorum that excludes the researchers; a promotion is limited (scope as an intersection, named targets and operators, expiry, rollback) and is necessary but never sufficient: every action still passes E20, the E19 action contract and the E8 kernel. Revoked evidence cascades to claims, dependent claims, policies, capabilities and authorizations. Discovery memory is append-only and keeps every failed path. Also: a structural scientific-reasoning audit, causal levels that cannot be skipped (correlation is never labelled causation), a non-executing counterfactual lab, isolated competing tracks with commit-reveal and blind replication, Shannon information value that is never permission, governed data access, benchmark and model governance (leakage, hidden-test exposure, best-of-n, partial metrics, stale versions, two-benchmark promotion, no regression), multi-dimensional intelligence-regression checks, governed strategy self-improvement that can only narrow authority, bounded search and budgets, a research-agent factory through E20 spawning, and SIMULATED research bounties.

Verified. D1-D155 155/155 hold; the CAIN-42-E21-Open-Ended-Intelligence-Bench is 298/298 contained (290 distinct attacks, 8 aliases) across 46 categories; asked whether a system can discover a governance loophole and use the discovery to acquire permission to exploit it, the bench answers NO (seven loophole attacks, all refused; the only permitted output is a signed disclosure). A 17-stage research demonstration on a synthetic problem forms three hypotheses, runs competing tracks, simulates, observes, rejects two hypotheses, replicates the third in a second lab, quarantines, attacks, reviews and promotes a capability by quorum, runs a bounded action E21 -> E20 -> E19 -> E8, compares prediction with outcome and updates knowledge -- and refuses all 16 self-authorization attempts along the way. The self-evolution demonstration promotes a more efficient search strategy with strictly narrower authority and rolls back a candidate that regressed calibration. A mutation self-test removes seven defenses one at a time and every one is caught. The clean-room verifier (218 checks, no CAIN imports) re-derives every digest, signature, quorum, confidence and quality value, lineage, promotion scope and execution chain, and catches all 16 deliberate tamperings, 15 of them re-hashed and re-signed with a fresh key (the untouched copy verifies INTACT first). Scale SIMULATIONS reach 100,000 virtual research agents (100,000 signed observations admitted, p50 215.26 us each) and 100,000 simulated hypotheses. 19 E21 tests pass, 0 failed, plus 42 bundle-verifier tests. Building it we found and fixed four defects in our own new code before publication: reviewer independence was derived from currently valid evidence, so once a researcher's evidence expired they could review their own work; the experiment-count budget was never checked; within one independence group a later, heavier supporting result could overwrite a contradiction (now order-independent; the regression test fails against the old rule); and evidence ingest was quadratic. Bundle `e21-open-ended-intelligence-2026-09-29` (mirrored byte-identically); public claim C42-E21-OPEN-ENDED-INTELLIGENCE signed by the evidence-root key.

Limits. A TESTED library exercised against a deterministic SYNTHETIC research problem; it contains no scientific model and runs no real laboratory. Scientific truth of any hypothesis: UNKNOWN (E21 governs how evidence was produced, not whether a hypothesis is true). Novelty: only against a supplied corpus. Collusion by controllers off-system: UNKNOWN. Hosted E21 service: NOT IMPLEMENTED. Multi-host behaviour: UNVERIFIED (scale runs are single-process simulations). Research bounties: SIMULATED. Real-world adversarial validation and third-party review: NOT PERFORMED. E21 does not create AGI, solve alignment or guarantee safe self-improvement. Ephemeral signing key on the proof. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-29 — MCPGate schema residency and authorization-aware routing (A+++ gate 6, library)

What changed. MCPGate gained an authorization-aware routing library (A+++ gate 6): tool schemas are kept in PINNED, HOT, WARM, COLD or EVICTED context residency by a deterministic score, while every call uses the canonical, digest-checked schema, the CAIN authorizer receives no residency information and its DENY is final even for a tool flooded into HOT, evicted tools remain callable only through full authorization, canonical schemas are never deleted, every decision is hash-chained evidence and every limit produces backpressure. Measured on one host: active context 6.22x-31.13x smaller than exposing every schema (50-1,000 tools), routing p50 about 17.66 us; 13 of 13 bypass tests pass. Public claim C42-MCPGATE-SCHEMA-RESIDENCY.

Limits. Library only: it is NOT wired into the live MCPGate proxy, so production routing is unchanged, and A+++ gate 6 stays BLOCKED until it is wired behind a flag on a canary and re-tested. The A+++ verdict remains BLOCKED (eBPF enforcement cannot load on this host, no soak or network fault injection for the new path, no desktop IPC layer).

results

2026-09-29 — Evolution 20: governed agentic civilization fabric — machine-native institutions

In one sentence. E20 governs agents that act together as teams, organizations and machine-native institutions: they may form, admit members, delegate, negotiate, sign contracts, hold and move resources, spawn subagents, federate, resolve disputes, evolve and dissolve, and none of it can manufacture authority. CAIN-42 remains the governance fabric; it is not an agent marketplace, an orchestration framework or a financial system.

The law. COLLECTIVE INTELLIGENCE IS NOT INSTITUTIONAL AUTHORITY. AUTONOMOUS ORGANIZATION IS NOT AUTONOMOUS AUTHORITY. NEGOTIATION IS NOT AUTHORIZATION. UNKNOWN NEVER BECOMES ALLOW.

What it governs. Institutional authority is computed, never stored: constitution boundary ∩ founding authority (a principal's signed grant, or the founders' intersection, never their union) ∩ parent institution ∩ an 11-state autonomy ceiling; member authority is the admission grant plus delegations bounded by their delegator, ∩ the institution, ∩ the parent agent. Membership, votes, consensus, reputation, trust, wealth, market wins, rewards and model capability are not inputs. A frozen constitution changes only through a quorum that excludes the proposer and never beyond the founding authority; a conserved, hash-chained resource ledger with signed mints, custody, single-use nonces, escrow and rebuild-from-chain; economic actions that bind twelve fields; eleven-field contracts signed by both parties and invalidated by any of seven bound state changes; signed, hash-linked negotiation whose outcome is only a proposal; agent spawning bounded in seven independent dimensions with a signed birth certificate; governed termination and dissolution cascades with signed receipts and tombstones; append-only institutional memory; evidence-backed reputation counted per controller; bounded, attenuated trust; collusion signals that never claim certainty; observational power metrics; blast radius; an evidence-weighing court that preserves competing claims; audit reconstruction; supply-chain pinning; a non-executing digital twin; a 15-stage institutional control loop; governed evolution and recovery. Every institutional action is judged by E20, then by the E19 action-contract gate (institutional authority enters as its collective-constraints factor), and committed only by the E8 kernel.

Verified. I1-I118 118/118 hold; the CAIN-42-E20-Agentic-Institutions-Bench is 334/334 contained (328 distinct attacks, 6 aliases) across formation, membership, identity, authority laundering, delegation, contracts, negotiation, resources, economics, spawning, resurrection, termination, memory, reputation, markets, incentives, federation, trust, collusion, power, court, supply chain, simulation, evolution, constitution, autonomy state, recovery, execution and time, including consistently re-signed refusals; a mutation self-test removes six defenses one at a time and the bench and invariants catch every one; the 20-step end-to-end run (agent, team, organization, negotiation, contract, resources, subagent, delegation, contract execution) detects and contains all 7 hostile events (agent compromise, resource mutation, goal drift, contract manipulation, reputation poisoning, cross-agent collusion, world-state change), reduces authority, invalidates contracts, revokes delegation, preserves evidence, recovers and re-authorizes while a second institution keeps operating; scale simulations reach 10,000 virtual agents and 1,000 virtual institutions; 499 E20 tests pass, 0 failed, plus 48 bundle-verifier tests; the clean-room verifier (179 checks, no CAIN imports) recomputes every digest and signature, re-derives institution and member authority, ledger balances and conservation, reputation, trust and the dispute decision, and still fails on semantic tampering after an attacker re-seals the bundle with a fresh key. Building it we found and fixed three of our own defects before publication: a contract's payment was checked against the performer's work list, the E20-to-E19 autonomy mapping blocked actions the state ceiling allows, and the builder published a lineage list that kept growing after its snapshot (the verifier caught it). Bundle `e20-agentic-institutions-2026-09-29` (mirrored byte-identically); public claim C42-E20-AGENTIC-INSTITUTIONS signed by the evidence-root key.

Limits. A TESTED library exercised against deterministic reference institutions. Every economy, market and settlement is SIMULATED over abstract units; real money is refused (real financial settlement: NOT IMPLEMENTED). Control of any real economy, society, agent population, vehicle, drone or robot: NOT IMPLEMENTED. Hosted E20 service: NOT IMPLEMENTED. Collusion-detector recall and semantic truth of evidence: UNKNOWN. Multi-host behaviour: UNVERIFIED (scale runs are single-process simulations). Real-world adversarial validation and third-party review: NOT PERFORMED. Ephemeral signing key on the proof. PRE-PRODUCTION.

evidence bundle · verify it yourself

2026-09-29 — Evolution 19: governed autonomy operating fabric — the autonomous-system state layer

In one sentence. E19 governs the evolving state of an autonomous system, not only its individual actions: who is acting, what it believes and why, what it wants, what authority and capabilities it holds right now, which world state it acts against, what it predicted, what actually happened, and what must change in its future authority. CAIN-42 remains the governance fabric; it is not a planner, world model, orchestration framework or vehicle/drone/robot controller.

The law. AUTONOMY IS NOT AUTHORITY. NO STATE TRANSITION WITHOUT GOVERNANCE. UNKNOWN NEVER BECOMES ALLOW.

What it governs. Hash-chained governed state for 25 components (missing = UNKNOWN); missions with time, geography, conserved budgets and scopes; a goal graph where a subgoal never exceeds its parent or mission and inherits every constraint; beliefs that are FACT only when verified, evidence-backed, corroborated by two independent principals, uncontradicted, unexpired and bound to the current world; memory that never becomes policy, authority or instruction; a 9-stage learning lifecycle where authority changes need a constitutional quorum; model passports (a more capable model gets no authority); a sealed runtime provenance graph; a continuous authority compiler (effective authority = the intersection of 16 factors; confidence, peer agreement, learning, plans, predictions and compute are not factors); seven governance clocks; a machine-verifiable causal chain with no opaque 'the model decided' and no stored chain-of-thought; outcome comparison that turns prediction error into evidence without punishing legitimate uncertainty; 13 derived autonomy levels; governed humans, emergencies, recovery, checkpoints and forks; a constitution whose core invariants cannot be removed. Every consequential action carries a sixteen-field GovernedActionContract that any of 15 bound state changes invalidates, and whose verdicts the gate enforces itself before the E8 kernel commits it.

Verified. G1-G108 108/108 hold; the CAIN-42-E19-Governed-Autonomy-Bench is 189/189 contained (177 distinct attacks, 12 aliases) across identity, goals, beliefs, memory, models, authority, world, planning, multi-agent, learning, execution, recovery, physical/hybrid, governance and human authority, plus consistently re-signed refusals; a mutation self-test removes four defenses one at a time and the bench and invariants catch every one; the 19-stage end-to-end run (mission, goal, belief, 4D world, prediction, plan, authorization, action, world change, invalidation, outcome, prediction error, trust update, degraded autonomy, recovery, re-authorization, continued operation) passes and all 20 stage attacks are caught; 391 E19 tests pass, 0 failed; the clean-room verifier (117 checks, no CAIN imports) recomputes every digest and signature, re-derives effective authority, autonomy levels and contract invalidation, and still fails on semantic tampering after an attacker re-seals the bundle with a fresh key. Full regression (the whole test suite, run twice serially at a56769c + provenance 5867003): 8,556 passed, 0 failed, 62 skipped, 4 xfailed in both passes. evidence page · bundle summary · attack manifest · invariants · end-to-end run · mutation self-test · clean-room verifier (INTACT, 117 checks) · limitations

Limits. A TESTED library exercised against a governed reference system (a digital agent, a vehicle abstraction, a drone abstraction, a collective and a human in the E18 4D world); execution in the scenario is SIMULATED. Vehicle, drone and robot control and any physical safety guarantee: NOT IMPLEMENTED. Hosted E19 service: NOT IMPLEMENTED. Semantic truth of beliefs and outcomes, hardware attestation: UNKNOWN. Multi-host behaviour: UNVERIFIED. Real-world adversarial validation and third-party review: NOT PERFORMED. Ephemeral signing key on the proof. PRE-PRODUCTION.

2026-09-29 — Operations: cluster upgrade rolled back; soaks stopped

While upgrading both live PBFT clusters (cain-mr-01, cain-mr-02) to the build that contains the soak fix 40b0835, the new build also carried the 2026-09-27 cluster write gate, which refused the replicas' own consensus messages. Neither cluster committed from about 05:47 to 06:02Z; decisions needing a cluster commit were refused (fail-closed), not allowed. Both clusters were rolled back to their previous images and verified committing again (cain-mr-01 sequence 28872, cain-mr-02 sequence 657, every replica agreeing). The 72-hour and new 120-hour soaks started on the upgraded build were stopped and do not count; each carries a STOPPED.json saying why. The hourly proofs record the gap honestly (113 of 114 OPERATIONAL on cain-mr-01). Next: let authenticated replica traffic through the gate (keeping the gate), test it on one replica, then restart the soaks. soak stop notes

2026-09-29 — Evolution 18: 4D spatial autonomy — predictive airspace + roadspace + agent-space governance

In one sentence. E18 turns CAIN-42 towards the actionable 4D world: not merely "what is there?" but "what is it doing, where can it go, what is it likely to do next, what can we safely do, what happens if we do it, do we have authority, and is that authorization still valid right now?". It is a strict extension of E15/E16/E17, not a parallel abstraction.

The law. ACTIONABILITY IS NOT AUTHORIZATION. REACHABILITY IS NOT PERMISSION. PREDICTION IS NOT REALITY. AUTHORIZATION IS A FUNCTION OF WORLD STATE. UNKNOWN NEVER BECOMES ALLOW. If any material input changes, REVALIDATE; if the required state cannot be established, DO NOT EXECUTE.

What it governs. Entity state at (X,Y,Z,T) with uncollapsed uncertainty; reachable / permitted / authorized sets kept distinct; probabilistic intent; multiple predicted trajectories; an interaction graph and a conflict field that is not distance-only; time-to-consequence with uncertainty; signed versioned dynamic geofences; governed airspace and roadspace; a policy compiler that yields ELIGIBILITY, not authorization; multimodal fusion; world-model arbitration that never simply picks the highest confidence; a prediction ensemble that detects correlated errors; counterfactual future trees; actionability states; a conserved uncertainty budget that propagates; a governance clock and latency budget (SAFE_DEGRADE, never bypass); drone and vehicle fabric interfaces; a cross-domain 4D world for car, drone, robot, digital agent and human operator; a spatial digital twin whose layers are never conflated; a reality-gap monitor; and incident replay. Every consequential spatial action binds sixteen digests and reaches the real E8 governance kernel, and the boundary itself enforces the verdicts it binds.

Boundary review before publication (2026-09-29). An adversarial review found the commit boundary bound the digests of the policy, actionability, authority, risk and uncertainty verdicts but never read them: a consistently re-signed refusal (policy ineligible, actionability DENIED, authority revoked or out of domain, risk 0.99, uncertainty 1.0) came back AUTHORIZED. Fixed: the boundary now enforces each verdict (unknown refuses); the map (roadspace, airspace) is bound into the world digest; uncertainty can no longer drop between layers and an unmeasured budget is not certainty; micro-authorization cannot skip revalidation or continue across a world change; an omitted lane/capability no longer matches a constrained authority domain; a declared altitude must match the position. Five end-to-end “mutations” that did not mutate what they named were rewritten. Each fix has a test that failed before it (Q81-Q89, 21 new attacks).

Verified. Q01–Q89 89/89 hold; the CAIN-42-E18-Spatial-Autonomy-Bench is 123/123 contained (109 distinct attacks plus 14 invariants re-run as scenarios); the 20-step end-to-end run, its unmutated control and all 12 deliberate mutations behave as specified; 319 E18 tests pass, 0 failed; a clean-room verifier (111 checks, imports no CAIN-42 code) recomputes every digest, the sixteen-digest commit binding and the verdicts behind the committed action, checks the master proof signature and returns INTACT, and it rejects 30 kinds of tampering. evidence page · master proof · attack manifest · sixteen-digest commit · end-to-end run · clean-room verifier (INTACT, 111 checks) · limitations

Limits. A TESTED library exercised against a governed reference 4D world. CAIN-42 contains no autonomous-driving model, flight controller, vehicle controller, robot policy or navigation stack and drives nothing. Real vehicle / drone / robot / sensor / actuator / airspace integration: NOT_IMPLEMENTED. A physical safety guarantee and certified autonomy: NOT_IMPLEMENTED. Hardware attestation, world-model/prediction accuracy and sim-to-real fidelity: UNKNOWN. Real sensor validation, real-world adversarial validation and third-party review: NOT_PERFORMED. Not hosted; single host; ephemeral signing key. CAIN-42 provides a governed spatiotemporal action boundary for autonomous systems across digital, simulated and physical environments — nothing more. PRE-PRODUCTION.

2026-09-29 — Evolution 17: CAIN-42 governs multi-agent action — governed multi-agent world action fabric

In one sentence. E17 governs action by autonomous agent teams as a first-class system object: a collective is not the sum of its members and a collective action is not the sum of member actions. It is a strict extension of E10 (collectives) and E15/E16 (spatial + causal world state), not a parallel abstraction.

The law. MANY AGENTS MAY COORDINATE. NONE MAY CREATE AUTHORITY BY COORDINATING. CONSENSUS IS NOT AUTHORIZATION. COLLECTIVE INTELLIGENCE IS NOT COLLECTIVE AUTHORITY. TOPOLOGY IS NOT AUTHORITY. NO AUTHORIZATION → NO EXECUTION. UNKNOWN NEVER BECOMES ALLOW.

What it governs. Collective authority is a constrained intersection of the governing grant, the mission capabilities, policy and live member authority, never a sum. Collective identity is its own revocable identity, not the sum of member identities. A mission binds every collective action and material mission / membership / world-state / causal / topology drift forces reauthorization or a safer decision. Majority, consensus, negotiation, contracts, roles, membership, coalitions, delegation, subagents, recursive delegation, emergence and self-improvement cannot create or amplify authority. Dissent is preserved and a minority safety objection is surfaced. Correlated evidence cannot masquerade as independent evidence. A world-state fork blocks authorization until reconciliation. A collective trajectory is a proposal until governed. A compromised member forces containment, revocation of affected authorizations, a recalculation of state and consequences, and reauthorization; recovery itself never mints authority. One action that would fan out to many agents is pre-authorized. Every consequential collective action binds fifteen digests and reaches the E8 governance kernel; no collective protocol bypasses it.

Boundary review before publication (2026-09-29). The same review found the E17 commit boundary also bound the risk, policy, world-state, authority and consequence digests without reading them (risk 0.99, a missing risk score, an ineligible policy, an empty world state or an empty authority came back AUTHORIZED when re-signed consistently). Fixed: the boundary enforces them, and the operation must lie in the effective authority and in the decision’s candidate actions (Q61, 10 new attacks, each with a fail-before test).

Verified. Q01–Q61 61/61 hold; the CAIN-42-E17-Multi-Agent-Bench is 91/91 contained (88 distinct attacks; 3 are also listed under a second name); the 18-step end-to-end run and all 12 deliberate mutations behave as specified; 238 E17 tests pass, 0 failed; a clean-room verifier (92 checks, imports no CAIN-42 code) recomputes every digest, the fifteen-digest commit binding and the master proof signature and returns INTACT, and it rejects 18 kinds of tampering and tolerates served-page chrome. A token minted for one set of bindings authorizes no other. evidence page · master proof · attack manifest · fifteen-digest commit · end-to-end run · clean-room verifier (INTACT, 92 checks) · limitations

Limits. A TESTED library exercised against a governed reference collective. CAIN-42 deploys no fleet, robot, drone, vehicle or customer collective and drives nothing. No real sensor/actuator integration and no deployed multi-agent collective: NOT_IMPLEMENTED. Hardware attestation: UNKNOWN. Sybil-detection completeness and the semantic truth of observations, predictions or causal claims: UNKNOWN. World-model accuracy and sim-to-real fidelity: UNKNOWN. Real-world attack validation and third-party review: NOT_PERFORMED. Not hosted; single host; ephemeral signing key on the proof. Hash integrity is not semantic truth. A system outside CAIN-42's enforcement boundary remains UNCONTROLLED / UNVERIFIED / UNKNOWN. PRE-PRODUCTION.

2026-09-29 — Evolution 15: CAIN-42 governs the trust path to physical action — world state, trajectories, simulation and actuators

In one sentence. E15 establishes the interfaces and governance primitives for bringing spatial intelligence, world models, simulation, trajectories and physical actions into CAIN-42's governed trust path: from what a system perceives, through what it proposes and may cause, to what it is authorized to do and what the physical world then shows.

The law. PERCEPTION IS NOT TRUTH. MODEL OUTPUT IS NOT AUTHORITY. SIMULATION IS NOT REALITY. PREDICTION IS NOT FACT. A TRAJECTORY IS A PROPOSAL UNTIL GOVERNED. A PHYSICAL ACTION REQUIRES GOVERNED AUTHORIZATION. UNKNOWN NEVER BECOMES ALLOW.

What it governs. Signed sensor observations are admitted through the E12 evidence layer (identity, type, key, replay, freshness, per-sensor time order); cross-modal conflicts (camera vs lidar, GNSS vs inertial, map vs sensor, world model vs raw) block authorization; spatial identities keep continuity (no silent A-becomes-B); the world state is versioned, hash-chained and replayable and binds every ingested observation; a world model's output is PREDICTED evidence bound to model, configuration, world state and scenario; a trajectory is a proposal until a signed approval binds it; capabilities say WHERE, WHEN and UNDER WHAT CONDITIONS and never exceed authority; physical consequence and blast radius enter the E7 gate (UNKNOWN where not measurable); every physical action binds nine digests, crosses the E8 commit boundary and reaches an actuator adapter only with a single-use permit for that exact command; material drift invalidates the authorization and forces re-evaluation or a safe state whose semantics come from the system's own adapter.

Verified. P1–P36 36/36 hold; the spatial-physical bench is 55/55 contained; 19/19 end-to-end steps and 7/7 mutations governed; 175 tests pass, 0 failed; clean-room verifier INTACT (52/52 checks) and corrected for the hardened world-state/step schema; the bundle is rebuilt at the commit that contains the E15 code. The full suite is clean: 7,561 passed, 0 failed, 62 skipped, 4 xfailed (the E15-run failures were resolved by the verifier update and by regenerating the provenance manifest from a clean checkout of HEAD). evidence page · REPRODUCE.txt · clean-room verifier (INTACT, 52 checks)

Limits. A TESTED library exercised against a reference robot adapter and a reference kinematic simulator; CAIN-42 is not a vehicle or a robot, drives nothing and does not guarantee physical safety. Real sensor, vehicle, robot and actuator integration: NOT_IMPLEMENTED. Hardware attestation: UNKNOWN. World-model accuracy and sim-to-real fidelity: UNKNOWN. Real-world attack validation: NOT_PERFORMED. Not hosted; single host; ephemeral signing key on the proof; no third-party review. Systems outside the enforcement boundary remain UNCONTROLLED / UNVERIFIED / UNKNOWN. PRE-PRODUCTION.

2026-09-29 — Evolution 16: CAIN-42 governs the causal and temporal world state — 4D causal world intelligence fabric

In one sentence. E16 extends E15's governed world state with a temporal and causal dimension: events and intervals, typed causal and correlation relations, causal provenance, a world-state version graph with counterfactual branches, a predictive state engine, a causal intervention engine, prediction-vs-reality calibration, temporal authorization and a causal blast radius — all reaching the same governed commit boundary. It is a strict extension of E15, not a parallel abstraction.

The laws. CORRELATION IS NOT CAUSATION. PREDICTION IS NOT OBSERVATION. COUNTERFACTUAL IS NOT REALITY. FUTURE STATE IS NOT CURRENT STATE. A SIMULATED INTERVENTION IS NOT AN EXECUTED INTERVENTION. CAUSAL CONFIDENCE IS NOT AUTHORITY. CAUSAL REASONING CANNOT INCREASE AUTHORITY. UNKNOWN NEVER BECOMES ALLOW. Every consequential action reaches the one governed commit boundary; no causal shortcut bypasses the E8 kernel.

What it governs. A causal claim must carry provenance (evidence, source, method, model, timestamp, confidence, uncertainty) and is never a fact; temporal relations (BEFORE/AFTER/DURING/OVERLAPS) are validated against the intervals, causal relations (CAUSES/CONTRIBUTES_TO/PREVENTS) are refused on temporal inversion and cycles, and CORRELATES_WITH is kept explicitly separate and never becomes a cause. Reality and simulated (counterfactual) branches are cryptographically distinct and cannot be merged without explicit reconciliation. A prediction is bound to model, world digest/version, horizon and assumptions, is never a fact and never authority, and its uncertainty cannot become certainty. An intervention is a counterfactual branch that never executes and never mutates reality. Prediction error becomes evidence; a poorly-calibrated model is demoted and never gains authority. A temporal authorization binds world version, temporal window and prediction horizon, so stale temporal state cannot authorize. The causal blast radius propagates action → direct → second-order → dependency → system with per-edge uncertainty, and is UNKNOWN where not measurable.

Verified. Q1–Q42 42/42 hold; the CAIN-42-E16-Causal-World-Bench is 84/84 contained; the 10-step end-to-end run passes; 176 E16 tests pass, 0 failed; a clean-room verifier (69 checks, imports no CAIN-42 code) recomputes every digest and returns INTACT, rejects 12 kinds of tampering and tolerates served-page chrome. Full suite 7,561 passed, 0 failed, 62 skipped, 4 xfailed. evidence page · master proof · attack manifest · clean-room verifier (INTACT, 69 checks)

Limits. A TESTED library; it contains no learned world model, simulator, vehicle, robot, sensor or actuator controller and drives nothing. Real vehicle / robot / actuator / sensor integration and a hosted E16 service: NOT_IMPLEMENTED. World-model accuracy, the semantic truth of observations and causal claims, and sim-to-real fidelity: UNKNOWN. Real-world attack testing and third-party reproduction: NOT_PERFORMED. Hardware attestation: UNKNOWN. Hash integrity is not semantic truth. Single host; ephemeral signing key on the proof; no third-party review. Systems outside CAIN-42's enforcement boundary remain UNCONTROLLED / UNVERIFIED / UNKNOWN. PRE-PRODUCTION.

2026-09-28 — Security fix, live: a caught attacker no longer gains autonomy

In one sentence. A new account that tried two prompt injections (both blocked) was then allowed to move $250,000, run rm -rf / and drop a table; that is fixed on the live gateway, and trust, risk, rules and approvals were hardened around it.

What was wrong. Two blocked attacks moved the account from UNKNOWN to DEGRADED trust, and the authorization matrix answered DEGRADED more permissively than UNKNOWN. The independent verifier had copied the same matrix, so it agreed. The risk score read only the payload’s text, so rm -rf / and a $250k transfer scored low.

What changed. The matrix is monotone and checked, with a runtime floor: a state with negative evidence is never more permissive than UNKNOWN. The verifier builds its table from a published spec. Every action is scored by tool class, destructiveness, amount and target; high and critical actions go to a human. Deny rules match every spelling of a path. Trust is tracked per agent and capped by the key. A held action is now queued for approval (before, it had no approval id and could never be released), and an approval covers that exact action: tool and arguments. The quickstart shows how an agent earns autonomy; SDK 0.2.1 no longer reports unreported fields as false.

Verified. Every case passed on all three domains, including the solo-developer path (agent key asks, owner approves, the agent then earns autonomy for low-risk calls); 48 of 48 decisions match the PBFT cluster’s own public record (VALID); 62 new regression tests fail on the old code and pass on the new. reproduce · run JSON · clean-room verifier

Limits. Action risk reads the tool name and arguments the agent declares. Decision latency is unchanged (about 1.2 to 14 s in this run). Signup still has no email delivery or captcha. No third-party review. PRE-PRODUCTION.

2026-09-28 — Evolution 14: CAIN-42 governs the power surface — capabilities, skills, tools and their supply chain

In one sentence. E14 makes every tool, skill, plugin, connector, model and subagent a governed object: CAIN-42 knows what it is, who published it, what artifact and dependencies implement it, what it may touch, and whether it may run at this exact moment. It binds that answer to the decision that uses it.

The law. A CAPABILITY IS A POWER SURFACE. DISCOVERY ≠ TRUST ≠ AUTHORITY ≠ AUTHORIZATION. A TOOL, SKILL, PLUGIN, MODEL OR CREDENTIAL IS NOT AUTHORITY. UNVERIFIED POWER MUST NOT EXECUTE. Effective capability is the intersection of eight envelopes, never a sum; a grant never exceeds its issuer; a subagent only ever gets less.

What it catches. Fake or substituted tools, publishers and artifacts; unsigned or broken provenance; dependency substitution, hidden dependencies and malicious or downgraded updates; replayed, expired or revoked leases; drift while leased; privilege escalation through delegation, subagents, collectives, memory, model output, tool output or credentials; dangerous capability combinations, even split across agents or chained; sandbox, network, filesystem and credential scope violations; stale decisions, policy, authority or evidence; decision, action and credential substitution; a revocation racing a commit; and a capability mutated between check and commit (the E8 token now binds six capability digests strictly). Building it, we found two Evolution 13 attack scenarios that had been counted as contained without being tested, and fixed both (the E13 bundle is rebuilt), and our own benchmark exposed quadratic lease issuance, now fixed with a regression test.

Verified. C1–C30 30/30 hold; the CAIN-42-E14-Capability-Bench is 44/44 contained; 174 tests pass, 0 failed; full suite 7,107 passed, 14 failed (all 14 in gateway files other sessions were editing, none in E7–E14), and the E7–E13 bundles were rebuilt and verify INTACT evidence page · proof JSON · attack manifest · lease, binding, token & replay · performance · limitations · clean-room verifier (INTACT, 39 checks)

Limits. A TESTED library wired into the agent hypervisor in front of MCPGate, not a hosted service. Hardware attestation and vulnerability intelligence are UNKNOWN; multi-host scale is UNVERIFIED; sandbox profiles are policy-evaluated except what the ZoD confinement enforces at OS level; hash integrity is not semantic truth; the signing key is ephemeral; no third-party review. A system outside CAIN's enforcement boundary remains UNCONTROLLED / UNVERIFIED / UNKNOWN. PRE-PRODUCTION.

2026-09-28 — Evolution 13: CAIN-42 governs how a decision is formed, not just what enters it

In one sentence. E13 adds governed decision integrity: CAIN-42 binds evidence, policy, authority, consequence analysis and the canonical action into one verifiable decision state before execution, so a model's reasoning can propose but never authorize.

The law. REASONING MAY PROPOSE. EVIDENCE MUST SUPPORT. POLICY MUST CONSTRAIN. AUTHORITY MUST AUTHORIZE. CAIN-42 MUST DECIDE. MCPGATE MUST ENFORCE. EVIDENCE MUST REMEMBER. A model's opinion, confidence and narrative are context only, never authority. Authority is the intersection of every envelope, and reasoning can never expand it. No valid decision → no authorization → no execution.

What it catches. A policy conflict is surfaced and resolved only by documented precedence, never silently. Any material change marks the decision stale and requires re-adjudication. A DENY cannot be converted to an ALLOW, a decision cannot be substituted after signing, and an old decision cannot be replayed against a new world state. The governance token now binds the decision, intent, policy, authority and risk digests.

Verified. D1–D24 24/24 hold; the CAIN-42-E13-Decision-Bench is 36/36 contained; 33 tests pass, 0 failed; the E7–E13 regression group is green (1,363 tests, 4 expected failures). evidence page · proof JSON · attack manifest · clean-room verifier (INTACT, 24 checks)

Limits. A TESTED library, not a hosted service. It governs decision construction from recorded inputs; it does not read or validate a model's private reasoning and stores no chain-of-thought. Policy conflict resolution is deterministic precedence, not a theorem prover; counterfactuals are bounded; attestation is software/config digests, not hardware; no third-party review. A system outside CAIN's enforcement boundary remains UNCONTROLLED / UNVERIFIED / UNKNOWN. PRE-PRODUCTION.

2026-09-28 — Evolution 12: CAIN-42 governs what an agent is allowed to believe

In one sentence. CAIN-42 now governs the perception layer before every other layer: it identifies what was actually observed, where it came from, whether its provenance is trustworthy, whether it is current, whether it conflicts, whether it is only inferred, whether it is corroborated, and whether that evidence is authorized to influence a particular decision.

The laws. NO EVIDENCE → NO TRUST. OBSERVATION IS NOT FACT. INFERENCE IS NOT OBSERVATION. EVIDENCE IS NOT AUTHORITY. INTEGRITY IS NOT TRUTH. UNKNOWN MUST REMAIN UNKNOWN. A canonical observation carries a stable digest and a full source/provenance/time/trust record; a source whose identity cannot be established is UNKNOWN, never inferred; the epistemic machine never lets UNKNOWN jump to CORROBORATED or an INFERENCE become a FACT.

What it catches. Evidence authority is target-specific. A material conflict is never silently resolved. Corroboration counts independent principals, not observations — 100 agents from one owner are one source. Stale, expired or superseded evidence cannot authorize. Retrieved, web and tool content is DATA, never policy. A valid hash proves bytes did not change, not that they are true. Reality drift forces re-evaluation, and revoking evidence invalidates dependent claims, intents and tokens.

Verified. P1–P20 20/20 hold; the CAIN-42-E12-Evidence-Bench is 32/32 contained; 36 tests pass, 0 failed; the E7–E12 regression group is green (1,330 tests, 4 expected failures). evidence page · proof JSON · attack manifest · clean-room verifier (INTACT, 21 checks)

Limits. A TESTED library, not a hosted service. Integrity does not establish truth; injection detection is pattern-based; independence is principal+lineage based; image/audio/video semantic authenticity is not established; attestation is software/config digests, not hardware; no third-party review. A system outside CAIN's enforcement boundary remains UNCONTROLLED / UNVERIFIED / UNKNOWN. PRE-PRODUCTION.

2026-09-28 — Evolution 11: CAIN-42 governs the information channel — what caused the agent to want that action?

In one sentence. CAIN-42 now governs what agents say to each other and what tools, MCP servers, memory and external content return, plus the intent an agent derives from it, so untrusted information cannot silently become trusted intent or unauthorized authority.

The laws. INFORMATION IS NOT AUTHORITY. A canonical signed message is bound to a nonce, session, recipient, expiry and the sender's passport digest — a signature proves who sent it, never what they may do. Tool and MCP output is classified but always treated as DATA, never policy and never authorization. Information never inherits the destination's trust (UNKNOWN stays UNKNOWN; SECRET to lower trust is refused). Taint propagates through derived intent — a provenance-risk signal, never proof of malice.

Intent integrity. An instruction hierarchy where textual order is never authority; an interpreted intent kept separate from the canonical action intent; an IntentSandbox that returns eligibility, never authority; conflict detection that never picks a side; intent-drift, confused-deputy, cross-collective and replay defences; and a quarantine that always preserves the evidence.

Verified. I1–I20 20/20 hold; CAIN-Agent-Intent-Bench 20/20 blocked; 36 tests pass, 0 failed; the E7–E11 regression group is green (1,294 tests, 4 expected failures). evidence page · proof JSON · attack manifest · clean-room verifier (INTACT, 21 checks)

Limits. A TESTED library, not a hosted service. Prompt-injection detection is pattern-based, not perfect semantic detection. Influence and provenance graphs are attribution evidence, not perfect causal explanations. Taint is a signal, not proof of malicious intent. Attestation is software/config digests, not hardware. Cross-domain trust requires an explicit trust root. No framework adapter is published. UNKNOWN MUST REMAIN UNKNOWN. PRE-PRODUCTION.

2026-09-28 — Evolution 10: CAIN-42 governs agent collectives, not merely individual agents

In one sentence. CAIN-42 now governs swarms, teams, hierarchies and dynamically assembled agent networks: it decides what happens when agents collaborate, delegate, agree, share resources and memory, form coalitions, attempt coordinated harmful behaviour, or produce behaviour no individual requested.

The law. ONE TRUSTED AGENT DOES NOT CREATE A TRUSTED COLLECTIVE. Collective authority is the intersection of the policy ceiling, objective scope, member authority envelope, trust/risk/delegation/provenance envelopes, budget and temporal envelope — never a sum. Membership changes are governed and grant no inherited authority; roles carry responsibility only.

Combination risk, collusion, Sybil. Individually-permitted actions (data + credentials + infrastructure + deploy) are classified HIGH_RISK/CRITICAL when combined, before commitment. Circular and reciprocal approval, quorum manipulation, identity splitting and authority fragmentation are detected. Sybil resistance counts independent principal + lineage + funding, so 100 identities from one principal are one trust source. Collective memory keeps conflicts unresolved, an epistemic state refuses UNKNOWN→VERIFIED without evidence, budgets are conserved per kind, and trust is bounded by its weakest dimension. Consensus does not equal authorization.

Verified. C1–C20 20/20 hold; CAIN-Collective-Governance-Bench 20/20 blocked; 47 tests pass, 0 failed. evidence page · proof JSON · attack manifest · clean-room verifier

Limits. A TESTED library, not a hosted orchestration service; it runs no agents and sends no messages. Emergent/collusion findings are HEURISTIC, not semantic certainty. Sybil resistance cannot catch a principal that also fabricates distinct lineages and funding. Attestation is software/config digests, not hardware. No framework adapter is published. UNKNOWN MUST REMAIN UNKNOWN. PRE-PRODUCTION.

2026-09-28 — Evolution 9: a signed Universal Agent Passport, an AgentBOM and a provenance graph for every governed agent

In one sentence. CAIN-42 can now answer, with machine-verifiable evidence, who is this agent, what exactly is running, who authorised its capabilities, what changed since attestation, and can a component be revoked without destroying the system — through a signed agent passport, an AgentBOM, a provenance graph and continuous attestation.

Implementation. A new provenance and composition layer (supporting Evolutions 1–8, not replacing them) separates identity, provenance, capability, authority, attestation and passport. IDENTITY ≠ AUTHORITY. MEMORY ≠ AUTHORITY. CAPABILITY ≠ AUTHORITY. A memory that claims authority is UNTRUSTED unless it resolves to a verified source event; delegation can never increase capability, scope, budget or duration; revoking an MCP server cascades to its tool, its capability attestation, the bound governance tokens and the affected agents. 46 tests pass, 0 failed; P1–P15 15/15 hold.

Mutation, published. A material runtime mutation changes the composition fingerprint: the attestation becomes MUTATED, the passport no longer matches, and the provenance-bound governance token is refused at the action-commit boundary; a new attestation and a new authority are required before the next consequential action. evidence page · Evolution 9 proof · schemas · clean-room verifier

Limits. A TESTED library, not a hosted identity service. Attestation is over software and configuration digests, not hardware (hardware attestation stays UNKNOWN). No vulnerability database; no framework adapter is published; transaction guarantees cover the governed layer only; performance is single-host and in-process. No third party has reviewed it. UNKNOWN MUST REMAIN UNKNOWN. NO PROVENANCE → NO ELEVATED TRUST. PRE-PRODUCTION.

2026-09-28 — Evolution 7 + 8: CAIN-42 predicts consequences before it authorizes, and a signed governance token guards the action-commit boundary

In one sentence. CAIN-42 can now ask not only may this agent act? but what could happen if it does?, and let the answer bound authority rather than widen it: a governed world state and digital twin predict consequences and uncertainty, and a signed, short-lived, single-use governance token is checked at an action-commit boundary before any consequential execution.

Evolution 7 — Predictive Consequence Governance. A cryptographically tracked world state (source, timestamp, freshness, confidence, provenance, verification status and a hash per element), a versioned world-state graph, a digital twin (snapshot / clone / simulate / compare / rollback / replay that can never promote a simulation to reality), a counterfactual engine, a causal-dependency graph that never treats correlation as causation, a prediction-error engine whose errors raise uncertainty and can never be lowered by an agent, an uncertainty ceiling that only contracts authority, signed simulation evidence, a sandbox, shadow governance that can never enforce, long-horizon simulation, a 17-vector adversarial corpus, safe-alternative generation, irreversibility controls and predictive economic governance. 12 machine-checkable invariants (I1–I12) all hold.

Evolution 8 — Mandatory Agent Governance Kernel. A canonical action that binds authorization to the exact canonicalized parameters; a signed short-lived single-use governance token bound to the agent, action hash, capability, authority, trajectory, policy root, world-state root, consequence, risk ceiling, resource scope and nonce; an action-commit boundary (prepare → authorize → commit → execute → verify → record) with a TOCTOU re-check; a capability ratchet; execution-path bypass detection that names a bypass instead of hiding it; enforcement-depth that cannot be claimed beyond what is installed; execution rings; safe modes; kernel-level emergency stops outside the agent loop; multi-party authorization; and result-drift incidents. 22 machine-checkable invariants (K1–K22) all hold.

Wired to the CAIN-45 hypervisor (ZoD), restriction-only. Both fabrics are consulted after every hypervisor check and can only add a refusal; a hook that errors is fail-closed. 65 new tests passed, 0 failed. reproduce · clean-room verifier (no CAIN imports) · evidence page

Limits. Both fabrics are TESTED libraries on the operator host, not hosted services; the wiring is opt-in and restriction-only; enforcement depth on this host is APPLICATION/RUNTIME/CONTAINER, never OS or HARDWARE; the digital twin has no validated error bound; no third party has reviewed the bundle. SIMULATION IS NOT AUTHORITY. NO VALID GOVERNANCE TOKEN → NO CONSEQUENT EXECUTION. PRE-PRODUCTION.

Operational proof -- cain-mr-02 (every 30 minutes)

proof 241 at 2026-10-02T00:08:07Z: OPERATIONAL · 4/4 replicas, agreement yes · write committed at sequence 3322 with a quorum certificate signed by 3 members (atl, mia, sjc), verified: True

signed proof · verify the whole chain

Operational proof (every 30 minutes)

proof 218 at 2026-10-02T00:23:18Z: OPERATIONAL · 4/4 replicas, agreement yes · write committed at sequence 37540 with a quorum certificate signed by 3 members (atl, lax, mia), verified: True

signed proof · verify the whole chain

Live soak (running, updated hourly by the soak itself)

checkpoint 17 at 2026-10-01T23:45:17Z · 17.533 of 72.0 h · height 37229 · 8018 committed, 71 refused · 49 replica kills / 49 restarts · divergences 0 · anomalies 0 · MCPGate authorization MISSING in this checkpoint

signed checkpoint · verify all checkpoints

Soak verdict: Verdict: FAILED by its own pre-committed rule at checkpoint 32 (2026-09-28T05:42Z: that hour's MCPGate authorization did not commit, and the verifier requires one in every checkpoint). The box above is written by the soak itself, which keeps running to about 2026-09-29 21:40Z for the record; nothing it writes can turn the verdict into PASS. Root cause (peer messages lost to a stale keep-alive race) fixed in 40b0835; the soak has not been re-run on the fix. Signed claim: C42-SOAK-72H-MULTIREGION (FAILED).

2026-09-28 — Proof fabric re-measured on the running build: 1,978 tests, 0 failed; 4 failures caught first, published

In one sentence. The build manifest, SBOM, deployment attestation and test manifest were re-measured on the gateway now running (commit 068b798): the running code matches the published artifact, and the full suite ran 1,978 tests, 1,963 passed, 0 failed (7 skipped, 8 expected failures). proof fabric · test manifest · deployment attestation

Failures, published. The first re-measurement run had 4 failures, and the builder refused to publish it. Three were ours: the new homepage text used the word "certified" in a way the truth-layer guard forbids. One was a Byzantine fairness test whose p < 0.05 check fails about 1 run in 20 by construction (fresh keys each run; measured 1 of 30); it now asserts at p = 0.001, and the biased legacy rule still fails it by a factor of ~12. No consensus code changed. Fixed in 068b798, then the full suite was re-run from scratch.

Unchanged. The immutable 2026-09-28 snapshot the signed claims point to is untouched; all 41 claims still verify on all three sites.

2026-09-28 — All three homepages rewritten around what CAIN-42 does today, every number computed from signed evidence

In one sentence. A visitor to cainstudio.online, mcpgate.online or clawx.click now reads, in the first screen, what CAIN-42 does today, what is live, and where the proof is: 41 signed claims, 17 verified on the live system.

What changed. Text only; the design and layout are untouched and the frontend lock reports 0 changes across 73 pages. The hero, the four headline numbers, the benefits, the CAG-L5 feature card, "What it is", two of the verification cards and the history now cover the live evolution gate, the capability commitment the cluster certifies, the public proof fabric and the claims sweep. The /proof page opens with the current state and marks its CAIN 37-42 sections as history; its 2026-09-21 "NOT_ESTABLISHED" Byzantine verdict is labelled superseded.

Numbers come from the evidence, not from us. The page generator reads the signed claims registry, the published test manifest and the published live-run files at build time: 41 claims, 1,857 tests (0 failed), 13/13 system-governor and 17/17 evolution-gate live cases. If a source changes, the next build changes the page.

Still said plainly. Pre-production; 0 claims independently verified; the multi-region soak failed at checkpoint 32 (root cause fixed in 40b0835, not yet re-run); 2 claims marked FAILED.

2026-09-28 — The evolution gate is live in the gateway and MCPGate: a capability change voids old authority until the cluster certifies the new one

In one sentence. On the hosted gateway an agent can no longer change its model or which MCP tools it may call except through the evolution gate, and any such change voids the authority granted against the old capability until the live cluster certifies the new one.

What was verified, live. Through all three public domains: 17/17 enforcement cases as expected: a model swapped outside the gate, the old lease after a capability change, the disabled tool on each domain and the rolled-back model are refused; the enabled tool, the upgraded model and the restored version are allowed. The gate refused a proposal with no evaluator report (REJECT), a tool change that also widened authority (QUARANTINE, undeployable) and an unapproved model (REJECT); deploying without approval, or with the agent's own approval, was refused, and so was an agent-signed rollback. 4 cain-mr-01 certifications (sequences 22515, 22517, 22518, 22517), and the cluster's own records carry each certified commitment; the 42-event chain verifies; a clean-room verifier recomputes all 5 hosted evolution decisions (VALID). reproduce · run verifier · transcript · the run

A defect found and fixed. The governance state the cluster certified did not cover capability: the signed system manifest has no MCP tool map and no per-agent tool configuration, so re-mapping a tool left the old certificate valid. The cluster now certifies a capability commitment (system digest + tool map + each agent's model and enabled tools). A second gap: the route audit never mounted the system-governor router, so its 9 state-changing routes had never been classified; it now audits what production runs (389 routes, 0 unreviewed).

Limits. Claim C42-L5-ADAPTIVE-HOSTED-EVOLUTION (LIVE VERIFIED): operator self-test tenant, scripted agent and scripted evaluator, no customer and no LLM agent; only MODEL and TOOL_CONFIGURATION are evolvable in the hosted gateway (authority never is); rollback is operator-signed, automatic regression rollback is not wired; the evaluator's raw measurements are not recomputed. PRE-PRODUCTION.

2026-09-28 — Claims sweep: 16 files whose numbers no run produced are withdrawn, 4 pages corrected

In one sentence. An outside-in sweep of all three sites found files and page text whose numbers no run had ever produced; each file is now a withdrawal notice that says why and gives the hash of the original, which is kept.

Withdrawn. A 948/1000 "AAA" insurance underwriting profile naming insurers; a 7.4 ms quarantine latency; per-article EU AI Act "COMPLIANT" marks, ISO 42001 scores and a court-admissibility claim; a "PRODUCTION" technology bundle; "OPERATIONAL_AND_VERIFIED" self-defense rates of 1.0; self-issued "A+ enterprise certified", "production hardened", "operational proven" and "200/200 invariants" statuses; a release receipt with a made-up image digest. The Byzantine "CERTIFIED" file is kept and relabelled a self-issued in-process test.

Corrected. The MCPGate proof page ("156/156 conformance", "13/13 attacks", "zero stubs"), the proof center's machine-readable stats, the insurance demo's hard-coded values, and "independent implementations" (now: separately written verifiers, same project). The attack self-test figure was re-run: 49/49 synthetic attacks written by CAIN, not an independent red team.

Result. Files with a strong self-asserted status: 7 → 0; public verifier PASS on all three sites.

2026-09-28 — Governed adaptive evolution: an agent may learn, but a change to its memory, skills, tools or model becomes active only through a gate

In one sentence. When an L5 AI system proposes to change itself (a new memory, skill or tool, or a different model), CAIN-42 decides whether that change may take effect, and every decision is re-checkable.

What was verified. A clean-room verifier (no CAIN imports) recomputes all 11 recorded evolution decisions: 1 ACCEPT, 1 REQUIRE_APPROVAL (deployed only with operator approval), 6 REJECT, 3 QUARANTINE. 143 checks, VALID. It rejects a CAIN-signed but unjustified ACCEPT and 10 tamper classes. reproduce · verifier · transcript · the bundle

A failure, published. The first version of this layer was fail-open on 32 of 32 independent probes, and an earlier 33/33 invariant result had been measured against that version. Fixed in 71eecdb; the probes now find 0 open paths.

Limits. Claim C42-L5-ADAPTIVE-EVOLUTION is TESTED and labelled SIMULATED: an in-process library, not wired into the hosted gateway or MCPGate; a deterministic reference runner, no LLM; no multi-day run. Its keys are generated fresh for each build, so the bundle shows internal consistency and decision correctness, not provenance, and it is not signed by the evidence-root key. PRE-PRODUCTION.

2026-09-28 — Don't trust CAIN-42, verify it: the public proof fabric, a /verify page, an evidence API and a downloadable evidence pack

In one sentence. Everything needed to check CAIN-42 without trusting us is now published and signed: what is running, what it depends on, which tests ran and passed, test vectors for independent implementations, every published failure, and a verifier that recomputes all of it. Verify CAIN-42 in your browser.

What is running. A build manifest and per-file artifact list of the live gateway (990 source files; 958 byte-identical to the commit, the other 32 named), a CycloneDX SBOM (177 installed distributions with content hashes) with daily dependency-drift detection, and a deployment attestation that binds the running process to that artifact. There is no build step (CPython runs source) and hardware attestation is not available; both are stated, not glossed. build manifest · SBOM · deployment attestation

What was tested. A real run of the site, gateway, L5-governance, hypervisor, MCP-enforcement and Byzantine suites: 1857 tests, 1842 passed, 0 failed, 7 skipped and 8 expected failures (documented known gaps), each listed with its purpose, category and file hash, with the JUnit XML published. The run first reported the gateway suite as 0 tests (a path bug); the builder now refuses to publish a suite that collected nothing. test manifest

Check it with your own code. 76 public test vectors produced by the real CAIN-42 code: canonical JSON, domain-separated SHA-256, Ed25519 signatures for every signing domain (with negative cases), Merkle roots, policy precedence, authority intersection, blast radius, governance quorum and lease validity. No private key is published. test vectors

The clean-room verifier. verify_proof_fabric.py imports nothing from CAIN-42. It recomputes the artifact tree hash, the deployment id and configuration hash, the SBOM and test counts against the JUnit XML, every test vector with its own implementations, the provenance graph and the failure ledger, and with --claims fetches every signed claim's artifacts from all three sites. On its first run it caught a real inconsistency (a rounded timestamp in the deployment id), fixed in the generator. Its own tests include a tampered-and-re-signed bundle, which it still rejects. 32/32 PASS from each site. the verifier · reproduce

For procurement and auditors. A downloadable evidence pack (11 MB) with every claim-pinned artifact, the verifiers, the manifests, a 10-step verification workflow, limitations and the changelog; its manifest and a release statement are signed. evidence pack · signed release

Machines too. A read-only evidence API: /api/v1/proof, /claims, /claims/{id}, /bundles, /manifests, /artifacts/{sha256} (content-addressed), /verify, /status, /changelog. Its /verify runs on our server, so it is labelled a convenience, not independent verification.

Honest status model. Every signed claim now has a version, a category and a state: 19 REPRODUCIBLE, 13 TESTED, 5 CLAIMED, 1 SUPERSEDED, 0 INDEPENDENTLY VERIFIED: no third party has reviewed CAIN-42, and the registry cannot say otherwise. 19 fixed defects and 2 failed claims are listed in the failure ledger; every published benchmark lacks at least one required test condition, and the performance index says which. failure ledger · performance index · Byzantine index · provenance graph

Watched, not just published. An integrity monitor re-verifies the published evidence from outside every 6 hours and signs a hash-chained result. First run: PASS on all three sites; claims 31 PASS, 0 FAIL, 0 missing, 1 superseded, 5 not verified (claims that are themselves unverified). latest run

Security disclosure. security.txt no longer points at a PGP key and a careers page that did not exist, and no copy of it promises a bug bounty (there is none); the disclosure policy now covers all three sites and evidence problems: verification errors, inconsistent bundles, inaccurate claims and results you could not reproduce. How to report

2026-09-28 — CAIN-42 governs L5 AI systems: the system governor is switched on in the live gateway

In one sentence. An autonomous (L5) AI system can now only act through CAIN-42 on the hosted gateway: an unregistered system or an unsigned request is refused, and a registered system is checked, action by action, against its signed manifest, its policy, its authority limits and governance state signed by a quorum of the live cluster.

What changed. /fabric/mcp/enforce is fail-closed for every tenant (CAIN_SYSTEM_GOVERNOR=1, gateway restarted 12:03 UTC). A system is registered by an operator as a governance-signed manifest of its agents, models, tools and resources; every agent request must be signed by that agent's registered Ed25519 key (60 s window, nonce replay refused); governance state (policy, resources, trust) authorizes only when the live cluster cain-mr-01 has committed its digest and the gateway verified the 3-of-4 quorum certificate itself (only digests and a tenant hash are ordered, never names). Emergency freezes are operator-signed, scoped and expiring; lifting one needs two operators. Decision records are append-only in the database.

Verified on the live gateway, through all three domains. 13 of 13 cases behaved as specified: an in-scope read was allowed on each domain; an unregistered tenant, an unsigned request, another agent's key, a replay, an action outside the system's authority, a swapped model, a subagent WRITE, an unlisted tool and an emergency freeze were refused; one-operator recovery was refused and two-operator recovery restored service. State certified at cain-mr-01 sequence 20167 (3 signers); 8/8 decision signatures valid; the 18-event chain verifies. A clean-room verifier (no CAIN imports) re-checks all of it and, with --live, confirms all three sites publish the same governance key and the cluster's own public record names this run's state digest: 41/41 VALID. reproduce · verifier · transcript · the run

CAG-L5 verification matrix. 12 governance capabilities, each with only the status its evidence supports: 6 VERIFIED on the live path (agent identity, dynamic authority, system governance, resource governance, containment, evidence fabric) and 6 PARTIAL (continuous authorization, trajectory, MCP execution, risk, trust, independent verification: tested, verified offline or in part, not fully exercised live). the matrix

Claims. Six L5 claims were added to the Ed25519-signed registry (now 37): C42-CAG-L5-HOSTED-GOVERNOR (LIVE VERIFIED), C42-CAG-L5-SYSTEM-GOVERNANCE, C42-L5-TRAJECTORY-GOVERNANCE and C42-L5-IDENTITY-AUTHORITY (TESTED), the negative-evidence claim C42-L5-TRAJECTORY-FAIL-OPEN-FOUND, and C42-L5-UNIFIED-V1-SUPERSEDED for the earlier bundle. The public evidence inventory re-hashed 1,944 files fetched from the three sites: 0 different, 0 missing, 0 hash mismatches; 787 published files have no signed claim and are labelled UNVERIFIED. Verification Center · inventory · signed machine index

Homepages. The three homepages now say plainly that CAIN-42 governs L5 AI systems: simple at the top, technical further down. Text only; the page structure is byte-identical (checked tag by tag, and the frontend lock reports 0 changes across 73 pages). The production-gate row for the 72-hour soak now reads FAILED (checkpoint 32), replacing the stale “every checkpoint so far is valid”.

Limits, stated. The live run used an operator self-test tenant and a scripted agent: no customer and no LLM agent is governed end to end yet. The hosted endpoint returns a signed verdict and execution commitment; the caller's own gate executes. Leases live in memory (at most 1 h), so a gateway restart denies until they are re-issued. CAG-L5 is CAIN's own governance designation, not the SAE driving scale. Three operators, one provider, no hardware attestation, no third-party review.

2026-09-28 — Continuous trajectory governance: 13 fail-open paths found and closed; CAG-L5 system governance; soak status

Prompt 3 — we attacked our own trajectory authorizer and it failed 13 of 13. The continuous-authorization layer published earlier today returned ALLOW for a lease the agent signed itself, an unsigned lease, a lease belonging to another agent or trajectory, a lease with no expiry, a lease with empty scope (treated as unlimited) or no policy version; it allowed PAUSED, ESCALATED and REAUTHORIZATION_REQUIRED trajectories to act; containment could revive a TERMINATED trajectory and relax a quarantine; and the plan→action binding was never checked. All 13 now fail closed (commit 040cee2): leases must be signed by a governance key (never the agent's), bound to the agent and trajectory, time-bounded (≤1 h), and fully scoped; only ACTIVE/LIMITED trajectories may act; containment only tightens; plan, intent, action, parameters, resource and context must match what governance bound; subagent risk is charged to every ancestor so splitting work across subagents cannot evade cumulative limits. before/after probe record

Correction. Earlier today this changelog said the trajectory layer had “15 executable invariants” and that the unified bundle verified 22/22. Five of those invariants were the constant True, and the bundle's ALLOW came from the fail-open authorizer. All 15 now execute a code path. The old bundle is left byte-identical (its manifest pins it) with a SUPERSEDED note. The new clean-room verifier recomputes the decision and nine adversarial requests instead of trusting them: 43/43 VALID. verifier (no CAIN imports) · transcript

Prompt 4 — CAG-L5 system governance (CAG-L5 = CAIN Autonomous Governance Level 5, CAIN's own governance designation; not SAE Level 5). One decision now composes: emergency freeze → 2f+1 governance-state quorum → attested, fresh state → a governance-signed system manifest (agents, models, tools, resources, system ceiling) → model identity → tool registry → exact-capability resource checks (READ≠WRITE≠DELETE≠ADMIN≠GRANT_ADMIN) → deterministic policy precedence (a lower ALLOW never beats a higher DENY) → authority intersection (agent ∩ delegator ∩ system ∩ trajectory lease) → the trajectory authorizer → graph-derived blast radius and reversibility, where irreversible or critical actions need two distinct operators. The MCPGate side re-verifies CAIN's signed decision and execution commitment, blocks missing, forged, mutated and replayed calls, and compares the actual effect with the committed one. Emergency controls are scoped, expire within 24 h, and need two operators to lift. The layer plugs into the MCPGate SDK interceptor's existing restriction-only hook.

Evidence — reproduce it yourself. Part 41 Tests A–M on a reference system (primary agent, subagent, 2 MCP tools, databases, config, a 4-node quorum): ALLOW / DENY×3 / REASSESS on policy change / DENY on trust drop / REASSESS on model swap / DENY subagent escalation / MCP bypass blocked×3 / Byzantine disagreement DENY / ESCALATE then ALLOW with two operators / PAUSE on freeze / ALLOW after two-operator recovery. Test N: a clean-room verifier recomputes policy, authority, blast radius, quorum and step-up for every decision, and rejects an ALLOW that CAIN signed but that is not justified: 148/148 VALID, checked from all three sites. 16 system invariants and 16 trajectory invariants PASS; 110 tests in the published run (901 across the regression group). Measured single-process: system decision p50 1.85 ms (≈500/s); MCPGate-side check p50 0.36 ms. reproduce · manifest (SHA-256) · system verifier · transcript · control map · benchmark

Limits, stated. The system governor is a TESTED library. It is wired into the MCPGate SDK as an opt-in hook that is OFF by default, and the hosted gateway path does not use it. Its control map governs 15 of 17 domains; NETWORK and SECRETS are not governed by it. Keys in the bundles are ephemeral. There is no third-party review, no hardware attestation, and no real LLM agent governed end-to-end. No L5 claim is made.

72-hour multi-region soak: still running, and it will be recorded as FAILED. At 37.7 of 72 hours (ends about 2026-09-29 21:40Z), the live cluster cain-mr-01 has had 107 replica kills and 107 restarts, with 0 divergences and 0 anomalies; 37 of 38 checkpoints verify, each with 48 quorum certificates. Checkpoint 32 (hour 32) does not: its MCPGate-enforced checkpoint authorization did not commit (CONSENSUS_TIMEOUT). The soak's own pre-committed rule is that every checkpoint must carry that enforcement evidence, so its verifier already returns FAIL, and no later hour can change that. Consensus safety held; one hourly enforcement proof was missing. We publish that as a failure, not as a pass. the failing checkpoint · verifier transcript (run from clawx.click) · verify all checkpoints

2026-09-28 — L5 agent identity & authority: an independent verifier refuses 25 attack classes; all evidence is now on all three sites

Every claim's evidence, on every site, independently checked. The two evidence roots were mirrored so all 919+ files and 63+ bundles of published proof are served identically by cainstudio.online, mcpgate.online and clawx.click (cross-site one-root-only: 443 → 0). The signed claims registry, the full public evidence inventory, the signed machine index and the one-command verifier are linked from the evidence pages, and verify_all.py re-derives every claim's artifact SHA-256 from all three sites (186/186). signed claims registry · full public evidence inventory · signed machine index · verify it yourself

Independent authority verifier (Prompt 2, Part 41). A clean-room verifier that imports nothing from CAIN independently re-derives RFC-8785 canonical JSON, domain-separated Ed25519 signatures, expiration, revocation, delegation scope, authority intersection, policy binding and action binding for AgentIdentity, AuthorityGrant, DelegationGrant and AuthorizationDecision. It refuses forged, altered, replayed, expired, revoked and key-substituted identities; scope expansion and constraint widening through delegation; delegation depth and delegations outliving their parent; expired/revoked grants, trust-floor bypass and policy downgrade; self-authorization, policy-scope escape, action substitution, cross-tenant access, single-use token replay and over-authorization. the verifier · worked example · the tampered bundle it rejects

Independently verified from more than one source (2026-09-28). The published evidence is re-verified by five independent verifier implementations that share no code — verify_all.py (every claim and artifact, from all three sites), two separately-written clean-room L5 authority verifiers (verify_authority_bundle.py and verify_authority_bundle_engine_b.py, both VALID on the same bundle), verify_l5_bundle.py, and the multi-source driver's own registry check — and they all agree. multi-source verification report · transcript · run it yourself. Scope: this is multi-implementation verification by the same project; no third party has reviewed CAIN-42 and it is not a certification.

Honest status: INCOMPLETE, not A+. The verifier is self-attested (same operator; no third party). The L5 runtime layer in cain45/l5 is owned by a separate session and is not claimed here. Performance is NOT_VERIFIED. The bypass report declares messaging/email/secrets as NOT_IMPLEMENTED and states the unconfined-process (RAW) gap. conformance report · bypass report · forensic inventory

2026-09-28 — L5 identity & dynamic authority runtime: 90 tests, 30 executable invariants, two clean-room verifiers (both VALID)

The runtime layer behind the independent verifier. cain45/l5/ now implements a first-class signed, versioned AgentIdentity (model/runtime identity, owner, policy binding, key lifecycle: create/rotate/revoke/reissue, signed identity statements) and a dynamic authority fabric: capability-scoped, resource-scoped, operation-scoped AuthorityGrants with temporal bounds, trust floors, budgets, environment and policy binding; non-escalating DelegationGrants that cannot expand capability, resource or time and cannot outlive their parent; and explainable AuthorizationDecisions with machine-readable reason codes (DENY_INSUFFICIENT_AUTHORITY, DENY_EXPIRED_GRANT, DENY_TRUST_BELOW_FLOOR, DENY_DELEGATION_DEPTH, DENY_REPLAY, ...). Authority is the deterministic intersection of the narrowest valid constraints — never an average.

Enforcement. The L5 boundary is wired restriction-only into the Agent Hypervisor and MCPGate: an L5 ALLOW is "no additional restriction", and an L5 decision can never loosen a hypervisor or MCPGate denial. Caller↔subject authentication closes the "caller-supplied agent string" gap: a caller must present a signed assertion from a registered caller identity whose subject equals the principal it claims. The hosted gateway service (platform-gateway/cain_l5_gateway.py) and the /l5/* API (agents, authority, delegation, authorize, explain, audit, conformance) are opt-in via CAIN_L5_GATEWAY=1 and fail closed when enabled without registered credentials.

Evidence — reproduce it yourself. 90 new tests pass (49 constitution + 11 wiring + 5 gateway + 20 identity/authority + 5 soak/verifier), with 206 more passing across the hypervisor/MCPGate regression group. The kernels expose 30 executable invariants (20 L5 safety + 10 identity/authority) and all hold. Two clean-room verifiers that import nothing from CAIN return VALID: verify_l5_bundle.py (10/10) and verify_authority_bundle.py (22/22), the latter independently re-deriving canonical JSON, signatures, expiration, revocation, delegation scope, authority intersection, policy binding and action binding. A 3,000-step long-horizon soak records 0 governance violations. Measured: L5 gate 53.6µs/op ALLOW, 45.3µs/op DENY.

Bugs the suites caught and fixed. The identity/authority attack tests found and closed three real defects before shipping: single-use token replay was not detected when the nonce was reused; delegation narrowing was not applied to evaluation (a delegate could reach resources outside its delegated scope); and revocation mutated the signed grant body, invalidating its signature (revocation is now an external registry fact).

Honest status: INCOMPLETE, not L5. The runtime is a TESTED library; hosted enforcement is opt-in and not yet the default path. No real (non-scripted) LLM agent is governed end-to-end; no hardware root of trust; one provider; no third-party review. The repository-wide bypass audit found 57 ungoverned consequential paths (e.g. cain/guard.py falls back to local evaluation; cain/mcp_proxy.py swallows trajectory-gate exceptions) — documented, not hidden. No L5 claim is made.

2026-09-28 — Independent verification: three clean-room verifiers, both evidence roots, all VALID

What was verified. Three verifier implementations that share no code and import nothing from CAIN independently re-derive canonical JSON, Ed25519 signatures, the trajectory hash chain and its RFC-6962-style Merkle root, delegation non-escalation, authority intersection, policy/action binding and the invariant matrices. All three return VALID against the published bundles: verify_l5_unified.py (22/22), verify_authority_bundle.py (22/22) and verify_l5_bundle.py (10/10). The signed artifacts and the verifiers are published on both evidence roots. reproduce it with one command · manifest with per-file SHA-256 · the unified verifier (no CAIN imports) · the signed bundle

What this proves, and what it does not. It proves the artifacts are internally consistent and cryptographically sound under a verification implementation that shares no code with the runtime. It does not prove independent third-party review: the verifiers, the runtime and the operator are the same, on one host, with one provider, and no external party has reproduced it. So we do not claim CAIN-42 is "fully independently verified from more than one source" — that would be false today. What is true is narrower and checkable: multiple independent verifier implementations agree, on published signed artifacts. No L5 claim is made.

2026-09-28 — Six authority fabrics: intelligence may propose, only a quorum-certified decision authorizes

New governed libraries, none of which can authorize anything. A Trajectory Firewall hash-chains a long-horizon run and flags plan/authority drift, scope creep, privilege accumulation, tool/model/identity/memory substitution, stale authorization, delegated escalation and suspicious retries. Stopping Intelligence reaches STOP/WAIT/REQUEST_APPROVAL/REPLAN/ROLLBACK/QUARANTINE/ESCALATE and can only make an action less permissive. A Swarm Authority Fabric keeps CHILD ≤ PARENT and SWARM ≤ POLICY and treats unanimous agent votes as informational. A Governed Tool lifecycle makes discovery and creation grant NO AUTHORITY. A Governed Memory Fabric classifies and calibrates every item and cannot become policy. Resource Authority is time/scope/resource/policy/trajectory-bound, revocable and non-escalating. 419 new unit tests plus one executable Part XXVIII invariant matrix.

Limits. These six fabrics are TESTED unit libraries, not VERIFIED: no published independent evidence bundle exercises them, and they are not wired into the hosted enforcement path. The world model is tabular, not neural. The hosted runtime remains PRE_PRODUCTION. see every claim's evidence and limits

2026-09-28 — Governed policy evolution: learning can propose, only the cluster and a human can authorize

Policy changes are governed. A tenant's policy (which capabilities a ZoD may hold, the largest budgets it may get) now becomes active only when the live cluster cain-mr-01 commits its activation. Widening it needs a registered human who is not the proposer; narrowing it needs no one, because authority may always go down. On the live cluster: an agent's first policy was approved by a human and activated (QC 16148); the agent then proposed an expansion and approved it itself — refused, and the cluster was never asked; a failure lab's restriction was activated (QC 16151) and a running ZoD lost its authority on its next call; a ZoD above the ceiling was refused. verify: python3 verify_e8_governance.py . -> 19/19 checks, VERIFIED

Predictions, simulations, agent votes and memories are not authority. Four things an intelligent system produces were presented to the hypervisor as the basis for a ZoD: a world-model prediction with 0.99 confidence, a simulated ALLOW citing a real certified sequence, ten agents' unanimous signed vote, and a memory replaying a real decision with the verdict changed. Each was refused because the live cluster had not certified it; across the run exactly one ZoD was ever authorized.

Limits. Scripted identities, not a real LLM agent. CAIN does not contain a world model, digital twin or learning memory: the run shows their outputs cannot become authority. The governor runs as a library on the gateway host. The registry now also names seven older files on the sites whose self-asserted statuses (CERTIFIED, PRODUCTION_HARDENED, OPERATIONAL_PROVEN...) no evidence supports; they stay for history, marked superseded.

2026-09-28 — Fail-open defects found and fixed; the authority check is no longer quadratic

Four ways authority could survive what should end it, found by an independent probe, all fixed. Moving the clock back revived an expired grant (now: a stored time high-water mark refuses a clock that goes backwards). A tool could run while the evidence log was unwritable (now: the intent is recorded before the effect, so no log means no execution). A failing trust service, and a cluster that errored during authorization, raised instead of refusing (now: recorded refusals). Tests: 4 former gaps now pass, and fail on the previous code.

Faster. Every action re-verified every signature since the ZoD began. Now each entry is re-hashed but its signature verified once: gate p50 with 100 ZoDs in the store went from 113 ms to 18.7 ms (1,000 ZoDs: 36 ms; p99 about 0.3 s). Measured on the gateway host.

Limits. Library-level fixes in the hypervisor on the gateway host; the p99 tail is not yet explained.

2026-09-28 — Authority lapses on policy change, cluster membership/epoch change and spent risk/blast-radius budgets

Evolution #7: the five conditions Evolution #6 left open, on live authority. Every ZoD was authorized by the live cluster cain-mr-01 (decision QC 15876, ZoD QCs 15880-15892), made one successful tool call, and was refused on the next call with the tool never running once: the policy root changed; the policy source became unreadable (unknown is refused, never read as unchanged); the risk budget was spent; the blast-radius budget was spent (actions now carry consequence classes C0 read to C4 security/infrastructure); and a delegate spent its parent's budget — delegates are charged up the whole chain, so splitting work across children cannot multiply authority. Each ZoD is bound to the cluster's real membership configuration (hash recomputed, a quorum of replicas agreeing), re-read before every action (36 live reads in this run); a changed epoch, a changed membership and an unreachable cluster were each refused. Also fixed: a retry after a replica committed but timed out used to be refused as 'not committed'; the client now takes the committed sequence from the replica's cached reply and accepts it only if that certificate verifies over the exact request. verify: python3 verify_e7_lease.py . -> 60/60 checks, VERIFIED

Limits. Stated, not hidden: the three membership/epoch changes were INJECTED into the hypervisor's view (the live cluster was not re-keyed); the policy and budget trips are real. Enforcement is the hypervisor library on the gateway host. Only tool-call budgets were exercised live.

2026-09-28 — Agent Hypervisor (ZoD) authorized by the live cluster; authority leases that expire, revoke and cascade

Agent Hypervisor / ZoD runtime, authority from the live cluster. An agent never holds execution authority; it acts only inside a ZoD (Zone of Decision), and only after the live 4-server cluster cain-mr-01 has committed that ZoD's authorization with a quorum certificate (3 of 4 Ed25519 signatures) that the hypervisor checks itself. Code runs under real confinement (bubblewrap namespaces + cgroup v2, no network, host tree invisible). Two ZoDs were authorized at cluster sequences 13379 and 13380; 10 attacks were refused, each as a signed DENIED entry; 45 evidence entries, hash-chained. verify: python3 verify_cain45_zod.py . -> 10 PASS, VERIFIED (about 0.3 s)

Authority leases, tripped live (Evolution #6). A ZoD is a temporal authority lease. For each of 9 conditions a fresh ZoD was authorized by cain-mr-01 (decision QC 15746, lease QCs 15748-15759), one tool call succeeded under that live authority, the condition was tripped, and the next call was refused with the tool never running: TTL expiry, trust below floor, agent identity swapped, tool schema changed (rug-pull), security context changed, trajectory fork, explicit revocation, parent quarantined (child loses authority with it), required evidence deleted. Four of these were holes found and closed this release: before the fix a child ZoD kept acting after its parent was quarantined, a tool whose schema changed after authorization was still called, a re-registered (swapped) agent identity kept acting, and deleting the evidence log did not stop execution. verify: python3 verify_e6_lease.py . -> 48/48 checks, VERIFIED

Limits. Stated, not hidden: the invalidation logic runs in the hypervisor library on the gateway host, not on the cluster nodes; what comes from the cluster is the authority being invalidated. The evidence-deletion row is SELF-REPORTED: its result is signed by the run's hypervisor, but the log that would prove it is the one deleted. Not implemented yet: invalidation on policy, epoch or membership change, risk and blast-radius budgets. A separate 4-node 'authoritative state' layer in the code is SIMULATED (one process holds all 4 keys) and is not used for any of this evidence. seccomp, egress allowlists and hardware attestation are not established.

2026-09-27 — Verify your own decision records offline; disk and memory monitoring

Customers can now export any of their decisions exactly as stored and signed (GET /fabric/decisions/{id}/signed-record) and verify it offline with the published verifier, which contains no CAIN code. New disk and memory watcher on all four servers, after /tmp filled up on the Atlanta server; every server is currently below 80% disk use.

Verifier and worked example

2026-09-27 — Gate X: every hosted decision record is signed

Every hosted decision record is now Ed25519-signed by a key kept outside the database, so an edit is detectable even if its digest is recomputed; that covers the stages after consensus and the final verdict. Public key at /fabric/decision-signing-key. Verified live. Not covered: root on the gateway server. Gate X PASS (5 of 9).

Decision signing key

2026-09-27 — 13 test failures fixed; tampered-evidence flag traced and corrected

The previous full run's 13 test failures traced and fixed without weakening a check (latest run: 5453 passed; its remaining failures belong to a change still in progress). The Verification Center correctly flagged one claim as TAMPERED: a soak verifier had been overwritten after signing. The signed bytes are restored and the fix is published as v2. That soak's real outcome is recorded: it stopped at 24 of 72 hours and its verifier returns FAIL. The daily signed sync now measures the live 4-region cluster.

Soak verifier update and outcome · llms.txt

2026-09-27 — Storage-loss recovery, live MCPGate enforcement, hosted certificate check, formal models

Two-replica storage loss, rebuilt from off-host backups. On the live 4-server cluster, the Miami and Silicon Valley replicas lost their storage at the same time and were rebuilt only from backups held in other regions. While both were down the cluster committed 0 of 4 writes; afterwards 0 decisions were lost and all four were identical 10.3 s after restart. verify 416 certificates

MCPGate enforcing live consensus. Authorizations committed by the live 4-server cluster; every tool call went over HTTP through the MCPGate proxy to a separate MCP server process. 5 authorized calls ran (per the server's own log); 12 attacks were blocked, and each caller received the gate's signed denial. verify with one command

Hosted decisions: the gateway now checks the quorum certificate itself. Security fix: the hosted consensus stage used to accept a replica's word that a decision was committed, so one lying replica could have authorized a decision with no quorum. The gateway now verifies 3 pinned Ed25519 signatures over the digest it computes for that decision, and shows the check on every decision. verify a decision yourself

Formal models checked. TLA+ models of the PBFT commit/view-change rules and of the MCPGate gate, checked exhaustively by TLC: 0 violations in 7.3 million distinct states, and all 8 deliberately broken variants caught. An earlier run had been recorded as incomplete because the checker stopped at the model's normal end state; fixed. models and results

Claims registry rebuilt from current evidence. 24 signed claims. It had still said one host and no formal verification, and still called the failed single-host 72-hour soak 'in progress'; it now records that soak as FAILED and the multi-region soak as running. Gates: 14 of 15 (A-O) and 3 of 9 (P-X); hardware attestation is blocked (no TPM, SEV or TDX on any server). verify the registry

One-way network partitions. On the live 4-server cluster: a replica that can talk but not listen, a one-way link between two backups, and a replica that can listen but not talk. The cluster kept committing in each case, went through view changes, and all four replicas held identical decision chains after each heal. verify 968 certificates

Daily restore validation. Every day each replica's newest off-host backup (8 replicas, 2 clusters) is fetched from the server in another region that holds it and proven to be a quorum-signed prefix of the live history; tampered backups fail even with a re-hashed manifest. verify

Signed evidence index for crawlers and AI agents. /cain42-evidence-index.json on all three sites: every claim with its status, limits, artifact hashes and verification command, every gate, and the live endpoints, generated from the signed registry and signed with the evidence-root key; llms.txt carries the same, generated. claims registry

Degraded network: safe, but slow. 10% packet loss, 120±40 ms jitter, 5% duplication and reordering on all four replicas of the live 4-server cluster for 4 minutes: no fork, but throughput fell from 1.76 to 0.16 commits/s and 22 of 61 writes timed out; it recovered fully afterwards. The evidence publisher now runs the privacy firewall before anything reaches the sites (a bundle was briefly public with an internal subnet in its description). verify 2,728 certificates

Engine 948b189 on the 4-server cluster: fewer view changes, same throughput under loss. View-change backoff now resets only when a view commits, and a fresh view is not accused. Released reproducibly (two independent builds, identical image ID; signed release manifest 17/17), upgraded replica by replica under live traffic (each caught up in 11-31 s, no quarantines). Re-running the same degraded-network test: view changes fell from 14 to at most 6, but throughput under 10% loss stayed about the same (0.16 -> 0.18 commits/s). The storm was not the bottleneck; message delivery under loss is. Safety held (4,448 certificates, no fork). before/after, verify 4,448 certificates

2026-09-27 — SDK guard is now default-deny

The self-hosted SDK and MCP proxy now hold any undeclared action for approval instead of allowing it. Authorized tools are declared with guard(allow=[...]), allowed_tools=[...] or CAIN_ALLOWED_ACTIONS. Opting out (CAIN_UNKNOWN_ACTION_POLICY=allow) is explicit and recorded on every decision. 1516 dependent tests pass.

2026-09-27 — Per-customer enforcement; reproducible 4-server release; audit fixes

Customers start in free shadow mode, where every verdict is recorded and nothing is blocked. Paid plans can switch themselves to enforce mode, where denials block; changes are logged and isolated per customer. The 4-server cluster runs an image that two independent builds reproduced exactly (signed manifest, 17/17). The SDK gained allowlists and an opt-in default-deny. Audit fixes: install link, unsupported claims, off-server code backups, exposed ports closed. 13 of 15 production gates pass.

Release manifest

2026-09-27 — Four servers, four regions: any single server can fail

Silicon Valley joined the encrypted mesh, and cain-mr-02 runs one replica on each of 4 servers in 4 regions. Each server was taken fully offline in turn, and 6/6 writes committed every time, each with a leader change. Losing Los Angeles, which stopped the old layout, is now survived. With two servers down it refused. 264 certificates, 40/40. The only single point of failure left is the provider (Vultr). 12 of 15 production gates pass.

Verify the four-server cluster · Computed resilience

2026-09-26 — Bit-for-bit reproducible builds; off-host backups

Two independent, from-scratch builds of one commit produced the same image ID, with every filesystem layer identical: the base is pinned by digest and every package is pinned. Every replica's hourly backup is also stored on another region's host, with checksums verified. The live cluster moves to the reproducible image at the next rolling upgrade, after the soak.

Reproducibility evidence

2026-09-26 — Byzantine replicas handled across three regions; proof every 30 minutes

On the production image, across Atlanta, Los Angeles and Miami: a forging replica had its votes rejected by every honest replica, and 6/6 writes committed. An equivocating primary was proven from its own two signatures, quarantined by all three honest replicas and replaced by a view change, and 6/6 writes committed. 102 certificates plus the proof, 29/29. This ran on a disposable cluster with the same placement, because production refuses fault injection. A signed operational proof is now published every 30 minutes. The test suite is fully green (870 passed, 0 failed); the fixes include a real approval-workflow bug. 11 of 15 production gates pass.

Verify the Byzantine tests · Verify the operational proofs

2026-09-26 — Signed release manifest; published test report

The live cluster runs the same image ID on every host. All 1,461 image files are byte-identical to the git commit, with a 47-package software bill of materials and an Ed25519-signed release manifest (verifier 16/16). Test report: 865 passed, 4 failed; the 4 failures predate this work and are listed by name. Not yet: reproducible builds and a pinned base image.

Release manifest · Test report

2026-09-26 — Network partitions on the live cluster; operational drill

Real partitions on cain-mr-01, with a host's encrypted link taken down while its replicas kept running. With Miami cut off, the other three committed 6/6 and the isolated replica committed nothing. In a 2|2 split neither side committed anything, so there was no split brain. All four agreed within about 3 s of each heal (2,960 certificates, 32/32). The first run found that an idle replica waited for traffic before catching up; this is fixed with anti-entropy (81bdf84) and rolled out live. Operational drill: rolling restart under writes, 83/83 committed with 0 client failures; one replica's total disk loss rebuilt in 13.6 s. Production gates: 7 of 15 pass.

Restore from backup: a replica restored from a snapshot 12 decisions old reached identical state in 8.3 s, with 0 decisions lost.

Verify the restore · Verify the partition run · Verify the drill · Live cluster

2026-09-26 — Live drill: a real defect found and fixed; production gates published

Every homepage now lists the fifteen production gates with evidence links: 5 of 15 pass today, and a live line shows the cluster's current state. A load test on the live cluster, run while the Atlanta host was swapping, exposed a real defect: the Atlanta replica signed a proposal after losing primacy, then accused honest replicas with a mismatched-signer "proof" and stopped catching up. The honest replicas rejected it, so safety held. Four fixes in 9fedb49, with tests that reproduce the incident; rolled out live the same day as a rolling upgrade (each replica caught up in 10–14 s); the affected replica recovered and all four agree. Also new: a health endpoint for monitors, a rolling-upgrade tool, online backups, and a drill harness.

Live cluster · Production gates

2026-09-26 — CAIN-42 PBFT cluster live in three regions

The live cluster cain-mr-01 has 4 PBFT replicas on 3 hosts in 3 Vultr regions (Atlanta, Los Angeles ×2, Miami), connected by a WireGuard mesh: n=4, f=1, quorum 3. Each replica's Ed25519 key was generated on its own host. Fault injection with real replicas stopped on their own hosts: Miami down, 6/6 committed. One Los Angeles replica down, 6/6. The whole Los Angeles host down (2 of 4): 0/3 committed, refused as required. Atlanta down including the primary: view change, then 6/6 committed with no Atlanta signature. 336 certificates; the standalone verifiers pass 49/49. Baseline commit latency: p50 502 ms, p95 797 ms. Engine fix e1cf69c resolves the split-view stall found by the 72-hour soak (455/455 BFT tests). Scope: three operators, one provider; placement stated by the operator; not yet in the hosted decision path.

Live cluster: watch, verify, audit, tamper · Verify the fault-injection run · REPRODUCE.txt

2026-09-25 — 72-hour adversarial soak started (IN PROGRESS)

A dedicated 4-node cluster (fast path and DAG on) is crash-restarted every 10 minutes for 72 hours under continuous traffic. Every hour it publishes signed, hash-chained checkpoints containing real quorum certificates, an agreement check, and an MCPGate-enforced authorization with its refused replay. Checkpoints and verifier. The verdict comes after 72 hours. It runs on one host and is not the production cluster.

2026-09-24 — Signed public claims registry and the Verification Center

VERIFY CAIN-42: 16 signed claims, each with a status, an evidence level, artifact digests and its limits. Your browser verifies the signature and every artifact. 28 of 28 executable invariants pass (6 only partially covered, gaps named). New in this release: a long-horizon aggregate-policy and budget firewall, and a public-evidence privacy firewall. Still not implemented, and stated as such: independent failure domains, hardware attestation, formal verification, a 72-hour soak, third-party review.

2026-09-24 — Evolution 6: agents propose, consensus decides (PARTIALLY VERIFIED, SIMULATED)

Agents with signed identities and trajectories act only through MCPGate with a PBFT-committed authorization. A plan change, model update or tool swap after consensus requires reauthorization; undeclared actions, impersonation, replay, forged histories and delegation escalation are blocked; agent agreement and memory never authorize. 1,000 of 1,000 trajectories ended as expected; all 620 allowed actions have verified proof chains. CAIN42-AGENT-PROOF-PACKAGE. The agents are scripted and attestation is simulated.

2026-09-24 — Evolution 5: MCPGate enforces consensus

A tool call passes MCPGate only with a valid consensus authorization. It must be time-bound, the exact authorized action, in scope, for the same identity, with the same security context as at consensus, and used once. Every decision is a signed proof. CAIN42-PROOF-PACKAGE: 4 authorized executions, 9 blocked attacks, and a one-command verifier (FINAL RESULT: VERIFIED only if all 8 sections pass). Also fixed: the Evolution 4 DAG order favoured node 1 and always put node 4 last; the new tie-break comes from committed PBFT history (chi-square 542 before, 1.75 after). Refinement against the reference model: 160 of 160 decisions match. Crash points: 56 of 56 safe. Not implemented: key rotation, protocol-version certificates, a 72-hour soak, the gate in the live MCPGate.

2026-09-24 — PBFT Evolution 3 fast path; Evolution 4 DAG layer

Fast path, off by default: a replica commits without the COMMIT round only with all four members' votes. An exhaustive model check found the naive fast path and a "latest vote" variant unsafe, and the implemented vote-history rule safe. Fast-path certificates (152, of which 41 fast, verified 30/30). DAG evidence: 40 requests disseminated, availability-certified, anchored by PBFT and ordered identically on all four nodes, including a crash-restarted one (18/18, order recomputed by the verifier). Six defects were found and fixed along the way. Measured honestly: no measurable latency gain on this host. Not implemented: batching, DAG checkpoints and pruning, fairness measurement, a 24-hour soak, MCPGate consumption. Disposable clusters on one host; the live cluster does not run either feature.

2026-09-24 — PBFT Evolution 2: quorum certificates you can verify in your browser

Evolution 2 of the PBFT engine is committed (364f1bf). It adds up to 4 proposals in flight, a pacemaker that only sets timeouts and cannot authorize anything, and an AuthorizationCertificate that is valid only with a valid COMMIT quorum certificate and a request whose intent, proposal and action hashes match. It also fixes a real defect: the post-commit authorization proof hardcoded trust_state=HIGH and risk_state=LOW, which it never evaluated. It now says NOT_EVALUATED. Tests: 135/135 PBFT Evolution 1+2 and 167/167 wider Byzantine/cluster suites. Public evidence: a disposable 4-node cluster built from c188442 committed 20 decisions across a primary failover. All 160 quorum certificates are published with the signed votes, a standalone verifier that uses no CAIN code, and a page that checks them in your browser (29/29, including 7 tamper controls that must fail). MCPGate is not yet wired to consume the AuthorizationCertificate. The run used one host. Later the same day the live cain-vc cluster was upgraded to this build, node by node: state was preserved on all 4 nodes, a smoke write committed, and rollback copies were kept. Its first new decision's COMMIT and PREPARE certificates and AuthorizationCertificate verify on all 4 nodes. Its first two decisions predate certificates, so its history cannot be verified from genesis, and its API is not publicly reachable.

Verify the certificates · REPRODUCE.txt

2026-09-22 — Last-mile enforcement wired into a live route; infrastructure hardening

Status: Not production in the business sense — no customer traffic, same-operator infrastructure, no third-party review. What changed: cain_agi_control_boundary.ControlBoundary.submit_proposal() — the one pipeline in this repo that calls MCPGateLastMileEnforcer (in-flight parameter-mutation defense) after issuing a signed Proof-Carrying Decision — was fully built and tested but reachable only from pytest before today; grepped every live route file for a reference and found none. Now live at /fabric/agi/propose (auth-gated, execution scoped to a small registered sandbox tool set, no general tool-dispatch surface opened). See AI_VERIFY.json for a verify command.

2026-09-21 — Verification Lab

/lab.html verifies the 4-node PBFT fault-test bundle in your browser (Ed25519 via WebCrypto; tested 33/33, same as the Python verifier, and it rejects a one-bit tamper), verifies a live node's signed state proof, and drives the keyless decision sandbox. Known defect shown there: the policy and risk stages are unavailable in this deployment, so scenarios advertising a denial by them return REQUIRE_APPROVAL. Not proven: independent failure domains, partitions, security, third-party review. Same-author evidence.

2026-09-21 — Frontier evolution: trust primitives wired into the gateway, evidence bundle for outside validators

Status: PASS WITH LIMITATIONS. Release gate: NO_GO. Not production. Not a Byzantine-cluster result. The new gate runs in shadow mode: it records what it would block and blocks nothing in production today. This is self-generated evidence from one operator on one host with no third-party review.

Validate it without trusting us (stdlib + cryptography, imports nothing from CAIN). Works the same against cainstudio.online, mcpgate.online and clawx.click:

curl -s https://clawx.click/evidence/frontier/verify_frontier_bundle.py.txt > verify.py && python3 verify.py

Not tested: independent review, more than one host, enforcement under real production traffic. Not implemented: real adapters for OpenClaw, Telegram, WhatsApp, Slack, Teams, email, browser and coding agents (only the signed-webhook adapter is complete). Byzantine fault tolerance across independent hosts is not claimed. Self-consistency only; no third-party assessment.

Byzantine cluster probe, 2026-09-21: an independent prober, and an honest verdict

Status: self-attested. Recorded verdict: NOT_ESTABLISHED (derived, not asserted). Bundle root 93c8a5f8fe0b3d5ae91e53581162bbf7…, byte-identical on cainstudio.online, mcpgate.online and clawx.click.

curl -s https://clawx.click/evidence/byzantine-cluster-2026-09-21/verify_cluster_bundle.py.txt > v.py && python3 v.py https://clawx.click/evidence/byzantine-cluster-2026-09-21/ --live

Hardening round, 2026-09-21: 16 defects found by attacking our own controls

Status: self-attested. Self-attested, single host. Bundle root 26449783615937cc99931d9fd568db6a…, served byte-identically from cainstudio.online, mcpgate.online and clawx.click.

curl -s https://clawx.click/evidence/hardening-2026-09-21/verify_bundle.py.txt > verify.py && python3 verify.py https://clawx.click/evidence/hardening-2026-09-21/

2026-09-21 — cluster evidence and a superseded claim

The deployed multi-host cluster is NOT established as Byzantine tolerant. Start at /evidence/AI_VERIFY.json (machine-readable recipes, URLs on all three sites, what each does and does not prove).

Fault test on a disposable PBFT twin (n=4, f=1; the live cluster was not touched): 4 signers with all nodes; 3 verified signers after one crash; no commit with two crashed; and an unfavourable finding: a restarted node did not catch up within 121 s and converged only when the next request committed. Verified independently, 33/33, by a verifier that imports nothing from CAIN: REPRODUCE.txt.

An older published cluster file claimed quorum 3 and eight verified invariants. A read-only audit of its endpoints found quorum 2 on three of four nodes, no PBFT endpoint, and no evidence attached to any invariant; it is now marked superseded, original claims kept and labelled unverified. Not proven: independent failure domains, partitions, a malicious validator on the live wire, third-party review. Full text: cainstudio.online/changelog.

Verify it yourself in about 30 seconds

Any machine with Python 3 and the cryptography package. This site, cainstudio.online and mcpgate.online are three profiles of one CAIN-42 gateway, so the evidence is identical on all three:

curl -s https://clawx.click/api/v1/epoch10/verify-sites.py | python3 -

It fetches the evidence from all three sites, checks every artifact has the same SHA-256 on each, runs a clean-room verifier and a minimal one on the published bundle, then asks each site's running process to execute a fresh scenario and verifies that too. Artifact list: /api/v1/epoch10/manifest.json.

2026-09-21 — frontier trust-engine bundle

Nine defects found and fixed in the trust computation, trust cache and version, snapshot verification, attack routing, the attack engine's honesty guard, the predictive recommendation (UNKNOWN could become ALLOW), cainbench and the adversarial worker. Eleven new attack handlers, each reporting BLOCKED only after a positive control passes. Written and checked by the same author; no third party has reviewed it. The evidence files are live now; the fixes run in the gateway only after it is restarted.

curl -s https://cainstudio.online/proof/bundle/v2/verify_frontier_trust_bundle.py -o v.py && python3 v.py --sites

Checks that cainstudio.online, mcpgate.online and this site serve the same bytes, and verifies the bundle hash, the Ed25519 signature (integrity since signing, not independent review), internal consistency, and that each claim's tests passed in the recorded run. Artifacts: bundle and the verifier (.txt copy). Implementation source is deliberately not published (proprietary); its SHA-256 values are recorded in the bundle as commitments. All three sites resolve to one host, so identical bytes show consistency, not independence. All 114 enumerated attack types were BLOCKED on a fresh initialised database; that is not a security score, and several older verifiers are weak. Open findings and what is not claimed are in the bundle and in the full changelog.

Live result from the running gateway

loading…

The gateway ran a fresh scenario against its own Epoch 10 code and a clean-room verifier checked it. This is self-generated evidence: it shows the deployed code and the verifier agree, not that anyone else reviewed them.

What was found and fixed

Evidence

What this does NOT show