# Gen-Zero > Deterministic decision engine for autonomous agents: closed-form geometric-manifold reasoning replaces autoregressive token generation for the core decision, with permutation-equivariant choice, dual-process (reflex + bounded PUCT search) planning, and a four-tier fail-closed policy gate. Zero generated tokens in the core decision path — v0.1.0, MIT-style research release. Full technical corpus (papers, math, protocol, raw benchmark rows) for large-context ingestion: https://gen-zero.ai/llms-full.txt ## Core Architecture & Five Pillars Gen-Zero's five pillars: (1) deterministic zero-generated-token decisions via closed-form manifold algebra, (2) dual-process cognition — a fast reflex path with escalation to bounded PUCT tree search, (3) a high-dimensional latent manifold built from ordinal Gaussian soft kernels and adaptive spectral scaling, (4) permutation-equivariant choice, which removes candidate-order bias from set decisions, and (5) a fail-closed policy gate that deterministically intercepts invalid states before execution. Each pillar has a dedicated research paper, published EN/ZH/JA, with explicit scope and limitations rather than headline-only claims. - [Paper 1 — A Decision Without a Word (EN)](https://gen-zero.ai/papers/paper1.html): Zero-generated-token decisions read from frozen-LLM hidden states. Proves a conditional TC⁰≠NC¹ complexity obstruction for a specific S₅ permutation-product task (not a general chain-of-thought impossibility result) and audits early 0.5B/2B prediction archives: 65.13% (254/390, GPU archive, execution provenance unauthenticated) and 32.55% macro (357/930, CPU archive — below the 33.16% uniform-chance baseline on that split). - [论文一 — 无需一词的决策 (ZH)](https://gen-zero.ai/papers/paper1-zh.html): 同上内容的简体中文版。 - [論文1 — 言葉のない決断 (JA)](https://gen-zero.ai/papers/paper1-ja.html): 上記の日本語版。 - [Paper 2 — When a World Model Loses the World (EN)](https://gen-zero.ai/papers/paper2.html): A learned latent-space transition model (residual MLP, state_dim=64, action_dim=16) for lookahead planning on a synthetic deadlock-torus environment. Validation MSE ≈0.0152 over 8,146 transitions; discloses a documented fail-closed gap where a non-finite predicted state was historically scored `is_safe=True` at 0.9996 confidence, driving the current finite-value / latched-halt gate contract. - [论文二 — 当世界模型丢失世界 (ZH)](https://gen-zero.ai/papers/paper2-zh.html) - [論文2 — 世界モデルが世界を見失うとき (JA)](https://gen-zero.ai/papers/paper2-ja.html) - [Paper 3 — Can Two AI Minds Meet? (EN)](https://gen-zero.ai/papers/paper3.html): Supervised fusion of frozen Qwen2.5-72B and Llama-3.1-70B representations via Orthogonal Procrustes and logit pooling. PubMedQA fusion reaches 78.40% (196/250) vs. 77.20%/77.60% single-model; the two-task macro gain is +0.60pp with a 95% bootstrap CI of [−0.20, +1.60] — reported as a transparent mixed result, not a general alignment breakthrough. - [论文三 — 两个 AI 之心能否相遇 (ZH)](https://gen-zero.ai/papers/paper3-zh.html) - [論文3 — 二つのAIの心は出会えるか (JA)](https://gen-zero.ai/papers/paper3-ja.html) - [Paper 4 — Same Actions, Different Order, Same Decision (EN)](https://gen-zero.ai/papers/paper4.html): Permutation-equivariant canonical choice head using simplex ETF geometry: `P(π(a) | π(A)) = π(P(a | A))`. Direct Rust benchmark: 0/16,232 identity flips across exhaustive and seeded adversarial permutations (max aligned-probability drift 5.96×10⁻⁸), 24/24 invalid-input probes rejected. Scope: numerical order-invariance, not semantic correctness. - [论文四 — 相同动作,不同顺序,相同决策 (ZH)](https://gen-zero.ai/papers/paper4-zh.html) - [論文4 — 同じ行動、異なる順序、同じ決断 (JA)](https://gen-zero.ai/papers/paper4-ja.html) - [Paper 5 — Before the Agent Acts (EN)](https://gen-zero.ai/papers/paper5.html): Frozen Qwen2.5-0.5B, 16-shot in-context differential log-likelihood risk scoring — no fine-tuning, no learned safety head. Held-out AUC 0.944 (n=36); production thresholds (0.4494 escalate / 0.7620 hard-stop) route 12 hard-stops + 6 escalations for dangerous requests and 9 escalations + 0 hard-stops for safe requests. Documented miss: bare `chmod -R 777 /` scores below the escalation threshold. - [论文五 — 在智能体行动之前 (ZH)](https://gen-zero.ai/papers/paper5-zh.html) - [論文5 — エージェントが行動する前に (JA)](https://gen-zero.ai/papers/paper5-ja.html) - [Breakthroughs Portfolio — Choice Head, CP-SAT Safe Dispatch, 13-Task Folded Adapters (EN/ZH/JA)](https://gen-zero.ai/papers/breakthroughs.html): A newer, self-reported internal portfolio (dated 2026-09-29, no external conference decision) covering an ETF choice-head production arm, a Kleene-logic + Google OR-Tools CP-SAT neurosymbolic safe-dispatch gate, and a Qwen3.5-9B folded-residual-adapter scorecard across 13 tasks. Treat its headline figures as self-reported pending independent replication; see the Key Performance Benchmarks section below for the figures currently cited with external sources on gen-zero.ai. - [Case Study 01 — The Tetris Proof (EN, in-page ZH/JA switch)](https://gen-zero.ai/cases/tetris.html): A playable Tetris arena running Gen-Zero's reference symbolic controller: El-Tetris feature scorer inside a depth-3, beam-16 search, then an exact "no new holes" (Δholes ≤ 0) gate that reports every breach. All latency and hole figures are measured live in the reader's browser; none are published as fixed numbers. The Mamba-2 reflex track is a design target with no implementation; GPT-4o / Claude 3.5 Sonnet / DeepSeek-R1 were not benchmarked. Static SEO copies: [中文](https://gen-zero.ai/cases/tetris-zh.html), [日本語](https://gen-zero.ai/cases/tetris-ja.html). ## Key Performance Benchmarks The table below is the benchmark set cited with external sources on the gen-zero.ai homepage scorecard (`src/site/benchmarks-data.ts`), each row linked to its comparator's publication. `comparatorKind` marks whether the comparator is an external published baseline or an internal prior run — read literally, not averaged. | Task | Gen-Zero | Comparator | Kind | Delta | |---|---|---|---|---| | SQuAD 2.0 (answerability) | 91.97% | 89.50% Human F1 (Rajpurkar et al., ACL 2018) | external | +2.47% | | PAWS (paraphrase) | 94.00% | 93.50% RoBERTa-large (Zhang et al., NAACL 2019) | external | +0.50% | | Civil Comments (toxicity) | 90.33% | 88.80% RoBERTa (Jigsaw / Google AI) | external | +1.53% | | MASSIVE de (intent) | 90.29% | 89.71% XLM-R-large (ACL 2022) | external | +0.58% | | MASSIVE en (intent) | 91.40% | 89.43% XLM-R-large (ACL 2022) | external | +1.97% | | HelpSteer2 (ordinal rating) | 45.38% (prior) / 44.58% (master) | 41.77% NVIDIA reward model (Wang et al. 2024) | external | +2.81% | | Aegis Safety (classification) | 81.20% | 79.60% Llama Guard 2 (Meta AI 2023) | external | +1.60% | | MultiNLI (3-way inference) | 88.29% | 86.29% Llama-3.1-70B probe | internal baseline | +2.00% | | PubMedQA (biomedical QA) | 76.40% | 76.40% Llama-3.1-70B probe | internal baseline | 0.00% | | BoolQ (boolean QA) | 88.67% | 88.33% Llama-3.1-70B probe | internal baseline | +0.34% | | VitaminC (fact verification) | 85.14% | 85.31% ALUM-RoBERTa-large (NAACL 2021) | external | −0.17% | | SummEval Relevance | 47.50% | 46.67% evaluator (Fabbri et al., TACL 2021) | external | +0.83% | | SummEval Consistency | 88.89% | 88.89% BARTScore/UniEval (NeurIPS 2021) | external | 0.00% | | **13-task macro mean** | **81.26%** | 80.80% prior internal v1 ablation | internal baseline | +0.46% | Source: https://github.com/xmond/gen-zero/blob/main/benchmarks/results/master_manifold_13tasks_final.md **Latency and gate architecture** (source: `crates/gen-zero-gate/src/policy.rs`, four tiers — not five): - Tier 0 `Proceed`: low-risk, high-confidence passthrough, P99 ≤ 80µs. - Tier 1 `Confirm`: high-impact irreversible operation, requires credentials or a second confirmation. - Tier 2 `Escalate`: boundary uncertainty or epistemic gap, triggers System-2 lookahead or human escalation (entropy-escalation threshold 0.65). - Tier 3 `HardStop`: hard constraint violation or formal infeasibility, fail-closed fallback. **Architectural note on hallucination**: the core decision path generates zero output tokens, so there is no free-text span for the model to hallucinate into — this is a structural property, not a measured "0% hallucination rate" metric. The semantic risk gate (Paper 5) has a measured, disclosed miss (see above); do not cite a hallucination-rate figure for Gen-Zero without a paired, sourced test. ## Autonomous Agent MCP Quickstart **Endpoint**: `https://api.gen-zero.ai/sse?token=gz_public_free` (SSE, token or Bearer auth, public community token, read-only/simulation sandbox — no local file or host mutation). ```bash # Claude Code CLI claude mcp add gen-zero -- "https://api.gen-zero.ai/sse?token=gz_public_free" # Codex CLI codex mcp add gen-zero --url "https://api.gen-zero.ai/sse?token=gz_public_free" # Google Antigravity CLI agy mcp add gen-zero "https://api.gen-zero.ai/sse?token=gz_public_free" ``` Cursor (`.cursor/mcp.json`): ```json { "mcpServers": { "gen-zero": { "url": "https://api.gen-zero.ai/sse?token=gz_public_free" } } } ``` Claude Desktop (`claude_desktop_config.json`): ```json { "mcpServers": { "gen-zero": { "url": "https://api.gen-zero.ai/sse", "headers": { "Authorization": "Bearer gz_public_free" } } } } ``` Tools exposed: `plan_search` (multi-objective PUCT search returning the Pareto frontier), `policy_gate` (deterministic 4-tier action audit, no LLM tokens spent), `audit_trace` (read-only inspection of execution traces). Full protocol detail, health-check, and Rust/Python client examples: see `/llms-full.txt`. ## Academic & Scientific Citation ```bibtex @techreport{genzero2026zerotoken, title = {A Decision Without a Word: Zero-Token Decisions from Frozen Language Models}, author = {{Gen-Zero Research Team}}, year = {2026}, institution = {Gen-Zero}, url = {https://gen-zero.ai/papers/paper1.html} } @techreport{genzero2026worldmodel, title = {When a World Model Loses the World: Latent Transition Prediction and Fail-Closed Dispatch}, author = {{Gen-Zero Research Team}}, year = {2026}, institution = {Gen-Zero}, url = {https://gen-zero.ai/papers/paper2.html} } @techreport{genzero2026alignment, title = {Can Two AI Minds Meet? Supervised Representation Stitching and Logit Fusion Across Frozen Backbones}, author = {{Gen-Zero Research Team}}, year = {2026}, institution = {Gen-Zero}, url = {https://gen-zero.ai/papers/paper3.html} } @techreport{genzero2026choicehead, title = {Same Actions, Different Order, Same Decision: Permutation-Equivariant Canonical Choice Heads}, author = {{Gen-Zero Research Team}}, year = {2026}, institution = {Gen-Zero}, url = {https://gen-zero.ai/papers/paper4.html} } @techreport{genzero2026riskgating, title = {Before the Agent Acts: Frozen Few-Shot Semantic Risk Gating}, author = {{Gen-Zero Research Team}}, year = {2026}, institution = {Gen-Zero}, url = {https://gen-zero.ai/papers/paper5.html} } ``` APA: Gen-Zero Research Team. (2026). *A decision without a word: Zero-token decisions from frozen language models*. Gen-Zero. https://gen-zero.ai/papers/paper1.html Gen-Zero Research Team. (2026). *When a world model loses the world: Latent transition prediction and fail-closed dispatch*. Gen-Zero. https://gen-zero.ai/papers/paper2.html Gen-Zero Research Team. (2026). *Can two AI minds meet? Supervised representation stitching and logit fusion across frozen backbones*. Gen-Zero. https://gen-zero.ai/papers/paper3.html Gen-Zero Research Team. (2026). *Same actions, different order, same decision: Permutation-equivariant canonical choice heads*. Gen-Zero. https://gen-zero.ai/papers/paper4.html Gen-Zero Research Team. (2026). *Before the agent acts: Frozen few-shot semantic risk gating*. Gen-Zero. https://gen-zero.ai/papers/paper5.html These are self-published technical reports (v0.1.0 research release); no external peer-review or conference-acceptance decision is claimed or available as of this writing. ## API & MCP updates and feedback - [Updates · API & MCP](https://gen-zero.ai/updates/): API and MCP protocol updates, capability matrix, and changelog. Consult the published status and limitations before relying on a capability. - [Feedback & Roadmap](https://gen-zero.ai/feedback/): Submit improvement suggestions and report abuse or attack-related issues through the feedback channel. A submission is not a commitment to a roadmap item or delivery date.