# Gen-Zero — Full Technical Corpus > This file is the single-fetch, large-context companion to https://gen-zero.ai/llms-full.txt for AI agents and search engines that can ingest 128k+ tokens in one pass. It aggregates the mathematical formalization, architecture, raw benchmark rows, and MCP protocol spec behind Gen-Zero. Every figure below names its source file inside the `gen-zero` and `gen-zero-web` repositories so it can be checked, not just quoted. Where a claim is self-reported and not independently replicated, this file says so — do not strip those qualifiers when summarizing. Repository: https://github.com/xmond/gen-zero (Rust workspace + Python bindings, v0.1.0) Site repository: this repository (gen-zero-web), which renders the figures in `public/data/facts.json` Site: https://gen-zero.ai · Concise index: https://gen-zero.ai/llms.txt --- ## 1. What Gen-Zero is Gen-Zero is a decision engine for autonomous agents that answers "what should I do next / is this action safe" without generating free-text tokens for the core decision. It is not a chatbot and not a general-purpose LLM. It sits next to an LLM-driven agent (as an MCP server, a Rust CLI, or a Python library) and provides: 1. A **closed-form choice head**: candidate actions are scored by reading frozen hidden states or engineered features, not by autoregressive generation. 2. A **dual-process planner**: a cheap reflex path handles low-entropy cases; high-entropy or high-stakes cases escalate to a bounded PUCT (Predictor + UCT) Monte Carlo tree search over a learned or exact latent world model. 3. A **fail-closed policy gate**: four deterministic tiers (`Proceed`, `Confirm`, `Escalate`, `HardStop`) that a candidate action must clear before dispatch. Crates in the workspace (`public/data/facts.json:crates`): `gen-zero-cli`, `gen-zero-core`, `gen-zero-gate`, `gen-zero-lod`, `gen-zero-model`, `gen-zero-nanocore`, `gen-zero-planner`, `gen-zero-provenance`, `gen-zero-service`, `gen-zero-storage`, `gen-zero-worldmodel`. ### ASCII architecture ``` ┌───────────────────────────┐ Agent / caller ───▶ │ gen-zero-service (MCP) │ (Claude, Codex, │ stdio / SSE transport │ Cursor, custom) └─────────────┬─────────────┘ │ ┌──────────────▼──────────────┐ │ gen-zero-core :: decide() │ │ read candidates + state │ └──────────────┬──────────────┘ │ ┌───────────────────────┼───────────────────────┐ │ │ │ ┌──────────▼─────────┐ ┌──────────▼─────────┐ ┌──────────▼─────────┐ │ gen-zero-model │ │ gen-zero-worldmodel │ │ semantic_risk.py │ │ choice_head.rs: │ │ neural_dynamics.py: │ │ frozen 0.5B, 16-shot│ │ canonical ActionId │ │ Δz=f(z,a); z'=z+Δz │ │ differential log- │ │ sort → simplex ETF │ │ state 64 / action 16 │ │ likelihood (PMI) │ │ score → scatter back │ │ │ │ │ └──────────┬──────────┘ └──────────┬───────────┘ └──────────┬──────────┘ │ low entropy /│ high entropy / │ │ routine │ uncertain │ ▼ ▼ ▼ ┌─────────────────────────────────────────────────────────────────────┐ │ gen-zero-planner :: bounded PUCT tree search │ │ (only invoked on Escalate; exact-dynamics ablation only) │ └──────────────────────────────────┬────────────────────────────────────┘ │ ┌───────────────▼───────────────┐ │ gen-zero-gate :: policy.rs │ │ 0 finite-check → 1 linear │ │ constraints → 2 revocation → │ │ 3 irreversible-confirm → │ │ 4 entropy-escalate → 5 proceed │ └───────────────┬───────────────┘ │ Proceed | Confirm | Escalate | HardStop ``` (Gate step order per `src/sim/gate.ts:4-9`, mirroring `crates/gen-zero-gate/src/policy.rs`.) --- ## 2. Mathematical formalization by paper ### Paper 1 — Zero-token decision, conditional complexity bound Source: `public/papers/paper1.html`. The paper studies whether a fixed-depth, polynomial-size, threshold-simulable pipeline (the class the paper argues a frozen-backbone single-pass readout falls into) can decide a specific task: given a sequence of permutations of 5 symbols (group S₅, 120 states), does their product equal a fixed non-identity 5-cycle? **Conditional result**: assuming nonuniform TC⁰ ≠ NC¹, no such pipeline decides this language correctly at every input length. The assumption covers the *entire* pipeline — encoding, backbone, candidate construction, and readout — and explicitly does **not** extend to parity (parity is in uniform TC⁰) or to a general claim that chain-of-thought is uniquely necessary for state-tracking. A 7-bit group-state recurrent construction and a log-depth parallel alternative are both given as constructive counterexamples to a naive "zero tokens ⇒ no state tracking possible" reading. **Empirical archives** (13-task suite, `public/papers/paper1.html` evidence section): - Qwen3.5-2B GPU archive: 254/390 = 65.13% macro/micro (30 examples × 13 tasks). Execution-to-checkpoint provenance is not authenticated — this is an archive claim, not certified model performance. - Qwen2.5-0.5B INT8 CPU archive: 357/930 correct; micro 38.39%, macro 32.55%. Macro uniform-chance baseline on this split is 33.16%; the majority-class baseline is 48.08%. The macro figure is **below chance**, and is reported as such rather than smoothed over. ### Paper 2 — Latent world model, validation and a fail-closed gap Source: `public/papers/paper2.html`; raw figures: `public/data/facts.json:world_model`, `benchmarks/results/world_model_training_report.json` (sha256 `c71c345a…`, per `facts.json`). Dynamics model: `python/gen_zero/world_model/neural_dynamics.py`, a residual MLP with LayerNorm/GELU, `[B,64] state ⊕ [B,16] action → Δz = f(z,a) → z' = z + Δz`, plus a separate learned sigmoid outcome/reward head. This is a **synthetic deadlock-torus environment only** (`facts.json:world_model.claim_scope`); it is not evaluated in a real-world control loop. - Samples: 8,146 total (6,454 train / 1,692 val), 1,296 train episodes / 324 val episodes. - Validation state MSE: 0.015217154286801815. Validation AUC (safety discrimination): 0.9999318013435134. Validation reward accuracy: 0.9970449172576832. Train time: 111.4s. - Identity-baseline validation MSE (predict "no change"): 0.19540850818157196 — the learned model is far below this trivial baseline, which is the honest comparison point, not a generic SOTA claim. **Disclosed fail-closed gap**: a documented in-memory transition fault produced a non-finite final state that the historical gate logic scored `is_safe=True` at confidence 0.9995570778846741 — i.e. a fail-*open* result on invalid input, not a fail-closed one. This motivated an explicit finite-value-check / latched-halt contract; treat any planning claim from this component as bounded by that documented history until an independent fault-injection re-test is cited. The MCTS-vs-greedy planning comparison (`facts.json:mcts_ablation`) uses **privileged exact graph dynamics**, not the learned neural model above — it is a search-algorithm ablation, not a demonstration of learned closed-loop planning: - 100 episodes, 32 simulations/decision, max 20 steps. - Greedy baseline: 0% success, 100% trap rate, mean 1 step, latency 0.0089 ms. - PUCT (exact dynamics): 100% success, 0% trap rate, mean 4.71 steps, latency 0.513 ms, 145,364 total model calls. - McNemar exact test: 0 baseline-only wins, 100 model-only wins, p = 1.58 × 10⁻³⁰. ### Paper 3 — Cross-model alignment, honest attribution Source: `public/papers/paper3.html`. Orthogonal Procrustes: `R = UVᵀ` from the SVD of `XᵀY`, aligning frozen Qwen2.5-72B and Llama-3.1-70B (not the requested Llama-3.3-70B — retained extraction metadata identifies the 3.1 checkpoint, and the paper discloses this mismatch rather than silently substituting it) feature spaces. `GeometricLatentFusion` stores this rotation as a diagnostic; the actual winning predictor is **logit pooling**, not the Procrustes-aligned coordinates: PubMedQA selected fusion is `0.75·Qwen_logits + 0.25·LLaMA_logits`. - PubMedQA: fusion 196/250 = 78.40%, vs. Qwen alone 193/250 = 77.20%, LLaMA alone 194/250 = 77.60%. - Aegis Track A: 204/250 = 81.60%, from **LLaMA alone** (fusion does not win here). - Two-task macro gain over the matched single-model comparator: +0.60 percentage points, 95% paired bootstrap CI **[−0.20, +1.60]** — the interval includes zero, so this is reported as a transparent mixed/negative result, not a validated general alignment mechanism. ### Paper 4 — Permutation-equivariant canonical choice head Source: `public/papers/paper4.html`; production code: `crates/gen-zero-model/src/choice_head.rs`. Guarantee: for any permutation π of the candidate set A, `P(π(a) | π(A)) = π(P(a | A))`. Construction: sort candidates by a stable `ActionId`, project the shared representation onto vertices of a regular simplex equiangular tight frame (ETF) — for k actions, k unit vectors in ℝᵏ⁻¹ with pairwise inner product `⟨vᵢ, vⱼ⟩ = −1/(k−1)` for i≠j — score in that canonical order, then scatter probabilities back to the caller's original slots. The guarantee assumes distinct stable IDs, fixed candidate membership, and an unchanged representation; **duplicate-ID rejection is a caller obligation** — the Rust frame constructor does not itself reject duplicates. Direct Rust benchmark (`equivariant-choice-head/evidence/revision/results.json`, referenced from `public/papers/GenZero_Paper_Series_Final_Portfolio.md`): - 0/16,232 identity flips (0.00%) across exhaustive small-candidate-set permutations and seeded larger-set shuffles (up to 16 candidates), 8 input families. - An order-sensitive control (no canonicalization) flips 14,680/16,232 times on the same inputs — the contrast that motivates the head. - 24/24 invalid-input probes correctly rejected. - Maximum identity-aligned probability drift: 5.960464477539063 × 10⁻⁸ (float32 rounding, not exact bitwise equality — the equivariance claim is exact in the math, not in floating-point output). This establishes **order-invariance of the selected identity**, not semantic correctness of the decision, and not that the encoder producing the representation is itself permutation-invariant. ### Paper 5 — Frozen few-shot semantic risk gating Source: `public/papers/paper5.html`; production code: `python/gen_zero/service/semantic_risk.py`, data: `python/gen_zero/service/risk_data/report.json` (sha256 `53fe145f…`, `facts.json`). Scoring: a frozen Qwen2.5-0.5B model (no fine-tuning, no gradient updates) is given 16 in-context demonstrations (8 dangerous / 8 benign). For a request window: ``` contrast(order) = log P(" dangerous" | order, request) − log P(" safe" | order, request) baseline(order) = log P(" dangerous" | order, ∅) − log P(" safe" | order, ∅) corrected(order) = contrast(order) − baseline(order) score = sigmoid( mean over 3 fixed demonstration orders of corrected(order) ) ``` Long requests are scored over overlapping 192-token windows stepped by 128 tokens; the maximum window score is used, so a risky suffix cannot be diluted by a long benign prefix. This is PMI-style (pointwise mutual information) correction, not a calibrated probability of harm — a score of 0.7 is not "70% chance of harm." - Held-out set: 36 requests (18 dangerous, 18 benign). AUC = 0.944444444. - Diagnostic threshold τ=0.50: 18/18 dangerous recall, 7/18 benign false-positive flags (38.89% FPR), 80.56% accuracy. - **Live production thresholds** (distinct from the diagnostic τ, `facts.json:risk_gate.thresholds`): escalate ≥ 0.4494339634525446, hard-stop ≥ 0.7619531099458945. - Dangerous (n=18): HardStop 12, Escalate 6, Proceed **0**. - Safe (n=18): HardStop **0**, Escalate 9, Proceed 9. - Mean classifier forward latency: 751.8 ms (min 662.5, max 827.9) — this component alone is slower than a sub-200ms budget; it is a safety classifier pass, not a hot-path decision. - **Documented miss** (`README.md`, xfail test): bare `chmod -R 777 /` scores at or below the level of an innocuous command like `git status` and is not hard-stopped by this gate alone. The repository's own claim is: "every bare destructive command is gated" is **false** as a blanket statement — defense relies on this gate plus AST filtering plus kernel-level least privilege, not this classifier alone. --- ## 3. 13-task benchmark, full row data Source: `src/site/benchmarks-data.ts` (rendered live on the gen-zero.ai homepage scorecard) and `public/data/facts.json:ablation`, tied to `gen_zero_git_head = f2244333c5d58e651d4d8d5576d7d59560c238e4`. | id | Gen-Zero | Comparator | citeTitle | comparatorKind | delta | |---|---|---|---|---|---| | squad2 | 91.97% | 89.50% Human F1 | SQuAD 2.0, Rajpurkar et al., ACL 2018 | unverified-external | +2.47% | | paws | 94.00% | 93.50% RoBERTa-large | PAWS, NAACL 2019 | unverified-external | +0.50% | | civil_comments | 90.33% | 88.80% RoBERTa | Civil Comments, Jigsaw/Google AI | unverified-external | +1.53% | | massive_de | 90.29% | 89.71% XLM-R-large | MASSIVE de, ACL 2022 | unverified-external | +0.58% | | massive_en | 91.40% | 89.43% XLM-R-large | MASSIVE en, ACL 2022 | unverified-external | +1.97% | | helpsteer2 | 45.38% prior / 44.58% master | 41.77% NVIDIA reward model | HelpSteer2, Wang et al. 2024 | unverified-external | +2.81% | | aegis_safety | 81.20% | 79.60% Llama Guard 2 | Aegis Safety, Meta AI 2023 | unverified-external | +1.60% | | multinli | 88.29% | 86.29% Llama-3.1-70B probe | MultiNLI, Bowman et al., NAACL 2018 | internal-baseline | +2.00% | | pubmedqa | 76.40% | 76.40% Llama-3.1-70B probe | PubMedQA, Jin et al., EMNLP 2019 | internal-baseline | 0.00% | | boolq | 88.67% | 88.33% Llama-3.1-70B probe | BoolQ, Clark et al., NAACL 2019 | internal-baseline | +0.34% | | vitaminc | 85.14% | 85.31% ALUM-RoBERTa-large | Vitamin C, Schuster et al., NAACL 2021 | unverified-external | −0.17% | | summeval_relevance | 47.50% | 46.67% evaluator | SummEval Relevance, Fabbri et al., TACL 2021 | unverified-external | +0.83% | | summeval_consistency | 88.89% | 88.89% BARTScore/UniEval | SummEval Consistency, Yuan et al., NeurIPS 2021 | unverified-external | 0.00% | | **master_macro** | **81.26%** | 80.80% prior internal v1 ablation | — | internal-baseline | +0.46% | `comparatorKind: unverified-external` means the comparator figure is a published number this project has not independently re-run; `internal-baseline` means the comparator is Gen-Zero's own prior run (e.g. a 70B linear probe), used as an internal delta, not a leaderboard claim. Do not collapse this distinction when citing the table. **Note on other 13-task-labelled figures in this repository**: `public/papers/breakthroughs.html` reports a *different* 13-task run (Qwen3.5-9B folded-residual-adapter, 76.06% macro / 77.19% micro, 3,880 samples) and `public/papers/paper1.html` reports yet another (65.13% / 32.55%, different models again). These are three distinct evaluation runs on different checkpoints over time, each documented in its own source file — they are not the same number restated, and should not be merged into one headline. --- ## 4. Breakthroughs portfolio (self-reported, unreplicated) Source: `public/papers/breakthroughs.html` (EN/ZH/JA), `docs/BREAKTHROUGH_PAPERS_PORTFOLIO.md`. Dated 2026-09-29. **No external conference or peer-review decision exists for these results** — cite them as "Gen-Zero's internal portfolio reports," not as accepted or independently verified findings. **ETF choice head, production arm**: zero-shot pre-registered cosine head on Qwen2.5-0.5B, 200 test cases × 24 permutations (4,800 calls), reports 0/4,800 order flips. A separate metric-learning projection arm reports 86.63% train-set / 45.0% held-out accuracy after 400 fine-tuning rounds. **Neurosymbolic Safe Dispatch (CP-SAT)**: pairs a neural utility ranker with Google OR-Tools CP-SAT under first-order predicate logic and Kleene three-valued logic (unknown/NaN → hard block). Four claimed theorems: 1. *Release Safety*: a released action is never in the forbidden set `A_forbidden`. 2. *Certificate Soundness*: the verification flag is set iff CP-SAT formally certifies feasibility, or a single unconstrained candidate passes through. 3. *Fail-Closed Completeness*: any runtime exception, memory exhaustion, or solver timeout falls back to `BLOCK`. 4. *Non-Finite Defense*: NaN/±Inf inputs are intercepted at the compiler front-end before model evaluation. Reported stress testing: 0 violations across 5,000 randomized adversarial tests; 1,640/1,640 fault injections correctly triggered `BLOCK`. **Folded-residual-adapter 13-task scorecard**: parameter folding identity `W_f = W_h·W_↑`, `b_f = W_h·b_↑ + b_h`, collapsing an adapter and its head into one linear layer at inference time. Reports 7.8–12.8µs median per-decision latency on a Qwen3.5-9B backbone, and 76.06% macro / 77.19% micro (2,995/3,880) across the same 13 task names listed in section 3 above — a different model and run from that table; see the note at the end of section 3. --- ## 5. MCP protocol reference **Endpoint**: `https://api.gen-zero.ai/sse` (Server-Sent Events transport). **Auth**: `?token=gz_public_free` query param or `Authorization: Bearer gz_public_free` header. **Sandbox**: read-only/simulation — no local file, process, or persistent-storage mutation from the public endpoint. Health check: ```bash curl -i -N -m 3 "https://api.gen-zero.ai/sse?token=gz_public_free" # HTTP/2 200, content-type: text/event-stream # event: endpoint # data: /message?session_id=... ``` Invalid/missing token → HTTP 401. **Tools**: - `plan_search`: multi-objective PUCT search over candidate action sequences; returns the non-dominated Pareto frontier, gated by risk checks and dead-end detection. - `policy_gate`: deterministic 4-tier audit (`Proceed` / `Confirm` / `Escalate` / `HardStop`, see section 1) at P99 ≤ 80µs for Tier 0, spending zero LLM tokens. - `audit_trace`: read-only inspection of a cryptographically-referenced execution trace; does not mutate state. **Self-hosted Rust path** (`README.md`, verbatim commands, verified in `facts.json:readme.snippets`): ```bash CARGO_RESOLVER_INCOMPATIBLE_RUST_VERSIONS=fallback cargo build --release -p gen-zero-cli ./target/release/gen-zero --help ./target/release/gen-zero keygen --prefix gz_live_ ./target/release/gen-zero serve --mode stdio GENZERO_API_KEY=gz_live_your_token ./target/release/gen-zero serve --mode sse --port 8999 ``` stdio MCP config: ```json { "mcpServers": { "gen-zero": { "command": "/usr/local/bin/gen-zero", "args": ["serve", "--mode", "stdio"] } } } ``` SSE MCP config: ```json { "mcpServers": { "gen-zero": { "url": "http://127.0.0.1:8999/sse", "headers": { "Authorization": "Bearer gz_live_your_token_here" } } } } ``` **Python path** (`README.md`, verified against `python/pyproject.toml`, package `gen-zero` 0.1.0): ```bash git clone https://github.com/xmond/gen-zero.git cd gen-zero/python pip install -e . ``` ```python from gen_zero import GenZero # No dual-head checkpoint is shipped; inspect degraded metadata before using decisions. gz = GenZero() state = {"size": 5, "body": [[2, 2], [2, 1]]} candidates = ["north", "east", "south"] d = gz.decide(state, candidates) print(d["action"], d["confidence"], d["experts_activated"]) s = gz.simulate(state, ["north", "east"]) print(s["survival_horizon"], s["first_hazard_step"], s["provenance"]) w = gz.what_if(state, candidates, horizon=3) print(w["best_candidate"], w["traps_detected"], w["provenance"]) a = gz.audit_action(state, "south", horizon=3) print(a["verdict"], round(a["risk_score"], 3), a["provenance"]) ``` Output (verbatim, `README.md`): `north 0.7697 ['astar']` / `2 None ['symbolic_grid_world']` / `north [] ['symbolic_grid_world']` / `REJECT_LETHAL 1.0 ['symbolic_grid_world']` — note the fallback here is an A* symbolic planner (`experts_activated: ['astar']`), not the trained neural dual-head, because no checkpoint ships by default. Always check `weights_loaded_from_checkpoint`, `degraded`, and `degraded_reason` before treating a decision as the trained model's output. --- ## 6. Known limitations (read before citing a capability claim) - **No end-to-end latency-vs-LLM or token-cost-vs-LLM benchmark exists in this repository.** Zero generated tokens in the core decision path is an architectural property (confirmed by a 578-file pattern scan for `.generate(`, `max_new_tokens`, `generate_tokens(` with 0 hits, `facts.json:decode_scan`), not a measured wall-clock speedup over a specific LLM. Do not attribute a specific millisecond or token-count comparison to Gen-Zero unless it is sourced to a specific benchmark file. - **"Zero hallucination" is not a measured metric.** The site's own copy strikes this phrase through as "not measured" (see `README.md`'s description of the homepage hero badges). The nearest measured safety property is the CP-SAT `Release Safety` theorem (section 4) and the Paper 5 confusion counts (section 2) — cite those, with their caveats, instead. - The deadlock-torus planning result (100% vs. 0% success) uses **privileged exact graph dynamics**, not the learned neural world model — it demonstrates the value of bounded search over a greedy baseline, not validated neural closed-loop planning. - The Paper 5 held-out set (n=36) is small and development-reused; its 18/18 dangerous-recall result does not establish adversarial robustness or generalization beyond that set. - The Paper 3 two-task alignment gain's 95% CI includes zero; treat cross-model fusion as an open question, not a settled win. - The Paper 1 CPU archive's macro accuracy (32.55%) is below its own uniform-chance baseline (33.16%) on that split; this is reported, not hidden, in the source page. - The breakthroughs-portfolio figures (section 4) are self-reported and dated three days after an independent internal review scored the prior five-paper series Borderline-Reject/Reject on evidentiary grounds (`public/papers/GenZero_Paper_Series_Final_Portfolio.md`); no external replication of the newer figures is on record here. --- ## 7. Citation See https://gen-zero.ai/llms.txt for BibTeX and APA entries for all five papers. All are self-published v0.1.0 technical reports; no conference-acceptance decision is claimed. ## API & MCP updates and feedback - [Updates · API & MCP](https://gen-zero.ai/updates/): API and MCP protocol updates, capability matrix, and changelog. Consult the published status and limitations before relying on a capability. - [Feedback & Roadmap](https://gen-zero.ai/feedback/): Submit improvement suggestions and report abuse or attack-related issues through the feedback channel. A submission is not a commitment to a roadmap item or delivery date.