Hands-on engineering hardening across the runtime and the three public sites. Every number below is copied from a result file that exists on disk; the source is named next to it.
Honest limits: no independent third-party audit; no hardware attestation; one provider and one operator; Phases 4 and 5 are TESTED libraries not wired into the hosted enforcement path. See LIMITATIONS.json.
| # | Name | Extends (existing moat / invariant) | Shipped |
|---|---|---|---|
| 1 | Zero-Stub & Truth Lock | governance invariant #30 (Zero-Mock / Zero-Stub); public truth layer | fake cryptographic placeholders removed -> fail closed (cain_trifecta_monopoly.py x3, cain/enclave.py)hardcoded QuorumSeal HMAC key -> env key or ephemeral per-process (cain_quorumseal.py x2)clawx homepage retained log corrected: shadow->enforce, 16->67 claims, FAILED soak no longer 'Running now'tests/test_zero_stub_guard.py added |
| 2 | Security Hardening | Moat 2 (security-context continuity) | SafeExpressionEvaluator: source eval-of-compiled-code replaced by a real AST interpreter (cain_control_loop.py x2)wildcard-origin CORS with allow_credentials=True disabled in 36 servicesunauthenticated /test-cain-private route removed (main.py x2)guard extended to fail on source eval and unsafe CORS repo-wide |
| 3 | Performance & Truth-of-Path | Moat 1/4 via the built governance data plane | governance data plane deployed live (CAIN_AUTH_SNAPSHOT=1; published_snapshots>=1, circuit_open false)operator tool scripts/cain42_dataplane/grant.py to grant the fast path (admin-gated)<15ms / sub-millisecond / 'Enforced in Silicon' claims corrected to the measured hosted figureRAM tmpfs /tmp exhaustion reclaimed; public health recovered to 200 with no 502s |
| 4 | Artifact Supply-Chain Governance | Moat 2 + Moat 4 | cain_artifact_attestation.py: publisher-signed manifests, rug-pull/version/capability/tool change detectionCompositionRiskEngine: order-independent composition risk naming the dangerous pathtests/test_artifact_attestation.py |
| 5 | Governed Memory at Write Time | Moat 1 + Moat 4 | cain_memory_governance.py: storage-time policy gate over the governed memory storetrust-aware retrieval (quarantine/expiry/trust/poisoning), independent promotion, purge-keeps-evidence, cross-tenant isolationtests/test_memory_governance.py |
| 22 | Framework Integrations + Advanced Long-Form Chatbot | Moat 2 (security-context continuity) | cainstudio.integrations.*: ten framework adapters using each SDK's native pre-execution hook (langgraph, openai_agents, crewai, llamaindex, pydantic_ai, autogen, google_adk, claude_agent_sdk, mcp, universal)one contract: a refusal means the tool body NEVER runs; on_block tell_agent returns the reason, raise propagates; available() reports importable adapters; per-framework pyproject extrasverified against real frameworks: sdk/python/tests/test_integrations.py 11 passed (LangChain, OpenAI Agents, MCP, Google ADK, universal, Claude Agent SDK) + 44 client testschatbot: stronger model, max_tokens 4000, effort high, TOP_K 10, passages 6000, 45s budget, prompt rewritten for 350-900+ word detailed answers; local cap 900/1500 -> 6000/9000chatbot bug fixed: _classify_intent generic-opener-vs-specific ordering (13/13 tests, was 11/13)launcher fades when not clicked within 60s (was 90s) and the ctw-fading CSS that was missing now exists |
| 21 | Public Enforcement Surfaces Rate-Limited (frontier + fabric status) | Moat 2 | /frontier/* (incl. the enforcement decision point /frontier/authorize) accepted unlimited unauthenticated requests (60/60 -> 200); now a router-level guard at 240/min per IP + global backstop/fabric/status (which probes live dependencies on every call) now guarded the same wayverified live: 300-request bursts -> {200:240, 429:60} on both; single requests unaffectedtests/test_frontier_and_status_limits.py |
| 20 | Dead Routers Mounted + False 'Mounted Live' Claim Corrected | Moat 2 + evidence integrity | routers/frontier_phase3.py (7 endpoints) and routers/break_glass.py (6 operator endpoints) were never imported by main.py -> every path 404'd while the changelog said 'mounted live'; both are now mounted and verified liveGET /fabric/frontier/status now 200; POST /fabric/frontier/a2a/authority returns granted:true; POST /api/v1/break-glass/grants (no key) is 403 operator-gated, not 404frontier API rate-limited as a public surface: 300-request burst -> {200:240, 429:60}tests/test_frontier_router_mounted.py incl. a guard that fails if any routers/*.py APIRouter is not imported by main.py |
| 19 | Limiter Contention Fix (no 500s) + Fail-Closed Store | Moat 2 | ratelimit._get_conn sets a busy timeout (default 5s, configurable) with WAL + synchronous=NORMAL so concurrent writers queue instead of raising 'database is locked'check_and_count/_checked fail CLOSED on any sqlite3.Error (a broken limiter store denies, never opens)same 800-request flood: before {200:612,500:24,TimeoutError:164} -> after {200:600,429:200}tests/test_ratelimit_contention.py; corrected a false positive in the Phase 17 SQL guard (PRAGMA busy_timeout=number) |
| 18 | Abuse-Resistance for Public Surfaces | Moat 2 | ratelimit.check_and_count adds a per-IP GLOBAL backstop across all buckets (a client can no longer stay under each endpoint limit while sweeping many)main.public_rate_limit guard + a router-level dependency on the public proof API (600/min per IP, 1200/min global)verified live: 700-request burst from one IP -> 599x200, 92x429; ordinary use unaffectedtests/test_public_rate_limiting.py |
| 17 | SQL Identifier Safety (last open code-level audit deduction) | Moat 4 (storage-layer provenance integrity) | cain_sql_identifier_safety.py: strict identifier validation, a validated Identifiers allow-list, and set_clause/insert_columns helpers that build parameterized SQL by constructiontests/test_sql_identifier_safety.py enumerates every dynamic-SQL module with a justification and fails on a new unreviewed site or a request-derived interpolationreview of the live sites: runtime_engine {field} validated against allowed_fields; cain_private updates built from internal literals; cain_schema table_name validated against CAIN_SCHEMA_TABLESthree false positives in the detector itself were fixed before the list was set to ground truth |
| 16 | Real-Store Decision-Record Integrity Proof | Moat 4 | scripts/cain42_public/build_record_integrity_proof.py opens the live evidence store READ-ONLY, samples real decision rows, re-derives each digest, checks the Ed25519 record signature and flips a real stage verdict to show detection; mirrored with a zero-dependency verifierverified on real data: 12/12 records re-derived intact, tamper detected, 1663 decisions in store (420 signed)fixed three generator defects: sparse-dict KeyError in the tamper vector, a parent-key mismatch, and non-identical mirrorstests/test_record_integrity_proof.py |
| 15 | Signed Governance Status Receipt (clean-room verifiable) | Moat 4 | scripts/cain42_public/build_governance_receipt.py collects live /fabric/status, /frontier/status, tool schemas, git commit and 3-site health, signs the whole receipt (Ed25519 over sha256(canonical JSON)), and mirrors receipt.json + zero-dependency verify_receipt.py + SHA256SUMS + index.html to both rootsverify_receipt.py recomputes the digest, re-checks the signature and confirms enforcing-list/chain consistency offline; both mirrors VERIFIEDcaptured at 2026-10-01T05:12:38Z: commit 3d8435f, dirty, mode enforce, 14 stages / 11 enforcing, 1 tool schema, all 3 sites 200tests/test_governance_receipt.py |
| 14 | Public No-Key Stage-Proof Bundle (clean-room verifiable) | Moat 4 | scripts/cain42_public/build_stage_proof_bundle.py emits stage-proof-2026-10-01 to both mirrors: stage_proof.json, zero-dependency verify_stage_proof.py, SHA256SUMS, index.html, countersignedthree records (allowed, egress_denied, toolargs_denied) with distinct digests and a tamper vector; the verifier recomputes digests, confirms distinctness, detects the tamper and re-checks the Ed25519 signaturefixed a real defect the verifier caught: sort_keys reordered stage keys after hashing; stages are now canonicalized before hashingtests/test_stage_proof_bundle.py |
| 13 | Restriction Stages As Signed Evidence + Live Tool-Schema Enforcement | Moat 4 | proof + tests that the live egress/toolargs verdicts are inside the Ed25519-signed decision digest, and that rewriting a recorded egress deny into allow is detectedoperator schema management GET/PUT /fabric/toolargs/schemas (admin-gated, atomic file write, registry hot-reload); a registered tool enforces, an unregistered tool stays observelive scenario POST /fabric/try?scenario=malformed-tool-args returns BLOCKED with blocked_by ['toolargs']tests/test_restriction_stages_are_signed.py, tests/test_tool_schema_admin.py |
| 12 | Automatic Log Self-Heal + Tool-Call Validity Live Stage | Moat 4 | automatic decision-log recovery at gate startup (runtime.get_gate): a non-empty log that fails verification is archived (bad line named, intact prefix + quarantined original preserved) and a fresh chain starts; missing/empty is a fresh chaintool-call argument validity wired as the live 'toolargs' Fabric stage (decision chain is now 14 stages; /fabric/status reports registered_tools)default observe for unregistered tools; a registered tool's args must match its schema and grounded values must be present in the CONTEXT the agent was given, never in the args themselvestests/test_toolargs_stage_and_log_recovery.py |
| 11 | Decision-Log Self-Heal + Grounded Tool-Call Validity | Moat 4 | DecisionLog refuses to extend an unverifiable log (fail closed); first_bad_line() names the exact break; recover() archives the bad log (naming the line, preserving the intact prefix) and starts a fresh chainfound and recovered a live wedge: the frontier gate was denying every action because decisions.jsonl failed verification at line 7448; intact prefix + quarantined original archived under data/frontier/archive/cain_tool_arg_validity.py: deterministic schema + grounding validator for tool-call arguments (missing required, wrong type/format, unknown parameters, ungrounded/hallucinated values)tests/test_tool_arg_validity.py |
| 9 | Authority-Preserving Inter-Agent Delegation | Moat 2 + Moat 3 | cain_agent_delegation.py: ordered authority levels (NONE |
| 8 | Egress Screening Wired Into The Live Decision Path | Moat 2 | TrustFabricControlPlane._egress_stage runs on every hosted decision; appears in /fabric/status enforcement.stages.egress and in the decision chainobserve mode when a tenant has no rules (recorded, not blocked); enforcing once any rule is configured; hard-denies loopback/link-local/private/cloud-metadata regardless of rulesmode pinning: a tenant may opt down to shadow but may not opt egress screening downtenant-key-gated management API GET/POST/DELETE /fabric/egress/rules (agent-forbidden)tests/test_fabric_egress_stage.py |
| 7 | Default-Deny Egress & Isolation | Moat 2 | cain_egress_control.py: default-deny destination allowlist (host/.domain/CIDR, ports, schemes, tenant/agent scoped)bypass guards (name rule never grants a bare IP; CIDR matches IP hosts only; lookalike suffix rejected) and hard denies (loopback/link-local/private/metadata/.internal before any rule)IsolationProfile: mandatory sandbox baseline with per-setting weakness reportingtests/test_egress_control.py |
| 6 | Agent Identity Interop | Moat 2 | cain_agent_identity.py: SPIFFE-style workload identity (SVID) with signed, time-bound, registration-bound verificationOAuth 2.0 token exchange (RFC 8693) with monotonic scope/lifetime attenuation; delegation 'act' chain; authenticated caller binding the verified subjecttests/test_agent_identity.py |
| Metric | Before (per-call PBFT) | After (quorum-committed snapshot) | Speed-up |
|---|---|---|---|
| CONSENSUS stage p50 | 2268.6 ms | 0.8263 ms | 2746× |
| CONSENSUS stage p99 | 3937.8 ms | 3.4686 ms | 1135× |
Source: CAIN42_GOVERNANCE_DATAPLANE_LIVE_RUN.json (real replicas; throwaway gateway DB). The public
POST /fabric/try probe mints a throwaway tenant per call, so it is never covered by a snapshot and stays ~2 s.
| Suite | Result |
|---|---|
tests/test_zero_stub_guard.py | 4 passed |
tests/test_artifact_attestation.py | 11 passed |
tests/test_memory_governance.py | 16 passed |
tests/test_agent_identity.py | 14 passed |
tests/test_egress_control.py | 17 passed |
tests/test_fabric_egress_stage.py | 6 passed |
tests/test_agent_delegation.py | 12 passed |
tests/test_frontier_a2a_authority_wiring.py | 7 passed |
tests/test_tool_arg_validity.py | 12 passed |
tests/test_toolargs_stage_and_log_recovery.py | 9 passed |
tests/test_restriction_stages_are_signed.py | 5 passed |
tests/test_tool_schema_admin.py | 5 passed |
tests/test_stage_proof_bundle.py | 6 passed |
tests/test_governance_receipt.py | 7 passed |
tests/test_record_integrity_proof.py | 7 passed |
tests/test_sql_identifier_safety.py | 9 passed |
tests/test_public_rate_limiting.py | 7 passed |
sdk/python/tests/test_integrations.py | 11 passed, 5 skipped (frameworks installing) |
platform-gateway/tests/test_expert_chatbot.py | 13 passed |
tests/test_ratelimit_contention.py | 5 passed |
tests/test_frontier_router_mounted.py | 5 passed |
tests/test_frontier_and_status_limits.py | 6 passed |
tests/test_cain42_public_truth_layer.py | 50 passed |
tests/test_cain42_pillar5_trifecta_monopoly.py + tests/test_cain_enclave_attestation.py | 24 passed |
tests/test_substrate_gateway_parity.py | 256 passed |
tests/test_governance_snapshot.py + _gateway + dataplane bundle verifier | 84 passed |
tests/test_cain_memory_adversarial.py + _sdk + cain45_memory + memory_poisoning_gauntlet | 63 passed |
combined hardening sweep | 361 passed in 4.57s |
sha256sum -c SHA256SUMSpython3 ../_publisher/verify_publisher.py . ../_publisher/phase1-5-hardening-2026-10-01.json