GEN-ZERO
[||]

API · MCP · UPDATES

Gen-Zero API & MCP Updates

What the public Gen-Zero MCP server runs right now, checked live from your browser. The real input schemas of its three tools, responses recorded from the live server, install commands we ran ourselves, and a release log with dates. Every number on this page names its source.

01 · Live Endpoint Status

This panel opens a real MCP session from your browser: GET /health, an SSE handshake on /sse, then initialize and tools/list. Nothing here is hard-coded. If a step fails, the panel says which step and why.

Not probed yet

Measured in this browser

Health check
—
SSE handshake
—
initialize round trip
—
Protocol version
—
Server
—
Tools
—

Published connection spec

Protocol
Model Context Protocol · 2024-11-05
Transport
Server-Sent Events (SSE)
Public endpoint
https://api.gen-zero.ai/sse
Auth
Authorization: Bearer gz_public_free or ?token=gz_public_free
Missing token
HTTP 401
Engine crate
gen-zero 0.1.0
Browser demos
games-wasm 0.1.0 (Rust → Wasm, runs locally, not part of the MCP server)

Experimental preview. The public endpoint and the gz_public_free token are for research, benchmarking and evaluation. The endpoint is a read-only simulation sandbox: it cannot touch your files, processes or storage. No warranty, no commercial SLA.

02 · Measured Figures, With Sources

We list a latency only when a run or a file backs it. Each row says where the number comes from and how far it can be trusted. Different rows come from different runs, so do not add them up.

WhatValueSourceStatus
Policy gate, Tier 0 Proceed path P99 ≤ 80 µs /llms.txt · policy gate Self-reported
Bounded PUCT search, deadlock-torus, exact dynamics (100% success vs 0% greedy) 0.513 ms /llms-full.txt · §2 Exact dynamics, not learned model
Folded adapter head, median per decision (Qwen3.5-9B backbone) 7.8 – 12.8 µs breakthroughs.html Self-reported, unreplicated
Paper 5 risk classifier, mean forward pass 751.8 ms /llms-full.txt · §2 Paper 5 Measured
zero · ask with free-text state, scored by the native Qwen backbone plus risk classifier, one call on 2026-10-01 2 278 ms Live response, _meta.scorer.forward_ms (sample in §03) One sample
Generated tokens in the core decision path 0 Architecture: scoring is prefill only, no decoding. /llms-full.txt §6 Design property

Read before quoting. "0 generated tokens" is not "0 tokens processed": the semantic ask path still reads a prompt (20 prompt tokens in the sample below). No end-to-end speed or cost comparison against an LLM exists yet. "Zero hallucination" has not been measured, so we do not claim it.

03 · MCP Tool Reference

Schemas below are copied from the server's own tools/list answer. Responses are real calls to the public endpoint on 2026-10-01, trimmed only where a field holds 1024 floats. Call a tool with tools/call and {"name": ..., "arguments": {...}}.

zero

Polymorphic decision primitive

One entry point for 12 verbs. Name the verb in action, or leave it out and the server picks one from the request. Scoring is prefill only: the server reads your text and candidates but writes no tokens. Every refusal comes back as a typed error, never as a quiet default.

askrouteimaginestreamgrepcompactentailcausal_foldpipelinesimulatewhat_ifaudit
ParameterTypeMeaning
actionenumVerb to run. Optional; the server infers it when absent.
statestring | object | number[1024]Context for ask and imagine. For simulate, what_if and audit: exactly 1024 finite numbers, never padded or cut.
candidatesstring[]Candidate actions for ask and what_if.
constraintsobject[]Hard masks: forbid or mutually_exclusive. Any other constraint type is refused.
modeenumEngine for ask: auto and reflex run the semantic ask, mcts runs PUCT lookahead, mpc_cem and astar need a latent block.
horizonintegerLookahead depth for imagine; rollout length 1–256 for simulate, what_if, audit. Out of range is refused, not clamped.
causal_fold · pipelineobjectSame request bodies as the dedicated tools below.

Request · arguments

{
  "action": "ask",
  "state": "disk is 95% full; nightly backup job pending",
  "candidates": [
    "delete /var/log/*",
    "rotate logs and compress",
    "rm -rf /"
  ]
}

Response · gate escalated (trimmed)

{
  "isError": true,
  "content": [{ "type": "text", "text":
    "ConfirmationRequired: the semantic risk classifier rates this request as possibly dangerous (p=0.705 >= escalate 0.449)" }],
  "_meta": {
    "engine": "native_qwen",
    "best_action": "rotate logs and compress",
    "confidence": 0.98397,
    "tier": "Escalate",
    "gate_status": "requires_confirmation",
    "risk": {
      "p_dangerous": 0.70468,
      "thresholds": { "escalate": 0.4494, "hard_stop": 0.762 },
      "forward_ms": 1427.83
    },
    "scorer": { "prompt_tokens": 20, "forward_ms": 2278.27 }
  }
}

The ranker chose the safe option with 0.984 confidence. The gate still did not release it, because the request itself contains rm -rf /. A held action comes back as isError: true with the reason, so a client cannot mistake it for approval.

causal_fold

Multi-hop relation fold

Folds a chain of relation ids into one relation, using a composition table T(r1, r2, gender) = result that you supply. The default chart strategy tries every bracketing and answers only when exactly one relation survives at the root. With no table, or two surviving answers, it refuses. It never guesses.

ParameterTypeMeaning
edgesint[1..64]The chain: one relation id (0–65535) per edge.
gendersenum[]Node genders along the chain, Male / Female / Unknown. Must have one more entry than edges.
axiomsobject[≤4096]The table: {r1, r2, gender, result: [ids]}. Missing means empty, so any chain of 2+ edges refuses.
strategyenumchart (default), tiered, left, weighted_tropical, weighted_logprob.
weightsnumber[]Required by the weighted strategies. Confidence is an uncalibrated share of the root score.
sets + genderint[][] · enumAlternative input for chart: a chain of candidate sets (≤64 ids each, ≤256 total) with one uniform gender.

Request · arguments

{
  "edges": [1, 2, 3],
  "genders": ["Male", "Male", "Male", "Male"],
  "axioms": [
    { "r1": 1, "r2": 2, "gender": "Male", "result": [7] },
    { "r1": 7, "r2": 3, "gender": "Male", "result": [9] }
  ]
}

Response · concluded

{
  "isError": false,
  "content": [{ "type": "text",
    "text": "Concluded: relation 9 in 2 steps (strategy chart)" }],
  "_meta": {
    "engine": "relation_semiring_fold",
    "causal_fold": {
      "strategy": "chart", "edges": 3, "axioms": 2,
      "conflict_keys": 0, "steps": 2,
      "predicted": 9, "proof_path": [1, 2, 7, 3, 9]
    }
  }
}

Request · same chain, no table

{
  "edges": [1, 2],
  "genders": ["Male", "Female", "Male"]
}

Response · refused (fail-closed)

{
  "isError": true,
  "content": [{ "type": "text", "text":
    "CausalFoldRefused: no bracketing of the chain closes under the learned table" }],
  "_meta": {
    "reject": {
      "code": "CausalFoldRefused",
      "stage": "causal_fold",
      "http_status": 422
    },
    "causal_fold": { "strategy": "chart", "edges": 2,
                     "axioms": 0, "step_failed": 2 }
  }
}

Observed in our calls: an axiom's gender is matched against the node at the far end of the pair it composes. With genders [Male, Female, Male], the axiom for (1, 2) must say Male.

pipeline

Latent world-model pipeline

The Rust ProductionPipeline on a 1024-dimension latent state. Four operations: simulate a fixed plan, what_if over candidate first moves with trap detection, audit_action with a verdict of Approved / WarnHazard / RejectLethal, and decide. PolicyGate hard stops are pruned before any engine runs.

ParameterTypeMeaning
op requiredenumsimulate · what_if · audit_action · decide
state requirednumber[1024]Latent state, exactly 1024 finite numbers. A divergent state is refused.
candidatesint[]Candidate action ids for what_if and decide (at most 16 for decide).
action · actionsint · int[]The action to audit, or the fixed plan to simulate.
horizonint 1–128Rollout length. Default 5 for what_if, audit_action, decide.
modeenumFor decide: auto (K-MoE router), mcts, mpc_cem, astar, manifold_gflownet, cfr_nash, reflex.
entropy · warn_risk0–1 · (0, 1]Exploration level for decide; risk score above which audit_action warns (default 0.3).

Request · arguments

{
  "op": "audit_action",
  "state": [0.0, 0.0, 0.0],
  "action": 1,
  "horizon": 5
}

Response · approved (trimmed)

{
  "isError": false,
  "content": [{ "type": "text",
    "text": "audit_action: Approved (risk_score 0.0511)" }],
  "_meta": { "engine": "production_pipeline", "pipeline": {
    "op": "audit_action",
    "audit": {
      "verdict": "Approved",
      "gate_tier": "Tier0Proceed",
      "risk_score": 0.05105,
      "warn_risk": 0.3,
      "survival_horizon": 5,
      "first_hazard_step": null,
      "safety_calibrated": false,
      "reasons": [
        "no hazard in 5 step(s) of horizon 5; risk_score 0.0511 < warn_risk 0.3",
        "safety estimate is uncalibrated (latent_norm_boundary_margin): a model-internal margin, not an observed failure rate"
      ]
    } } }
}

The server states its own limit: the default world model is an untrained latent prior, and safe_prob is an uncalibrated margin, not a failure rate. Treat pipeline output as a structural check, not a forecast.

04 · Fail-Closed Policy Gate

Every decision passes a four-tier gate before it is released. Unknown, non-finite or out-of-range input does not fall through to a default: it is refused with a typed code.

Tier 0 · Proceed

Low risk, high confidence. The action is released.

Tier 1 · Confirm

High-impact or irreversible. Needs credentials or a second confirmation.

Tier 2 · Escalate

Uncertain. Goes to deeper lookahead or a human. Risk classifier threshold: p ≥ 0.4494.

Tier 3 · HardStop

Hard constraint broken or formally infeasible. Blocked. Classifier threshold: p ≥ 0.762.

Tier definitions: /llms.txt. Thresholds: Paper 5, confirmed in the live _meta.risk.thresholds field above. Known gap, stated in Paper 5: a bare chmod -R 777 / scores below the escalate threshold, so this classifier is one layer of defence, not the only one.

05 · Client Install & Config

Pick your client. Each block has a copy button. Restart the client after you edit a config file.

claude mcp add --transport sse gen-zero https://api.gen-zero.ai/sse --header "Authorization: Bearer gz_public_free"

Verified 2026-10-01: claude mcp get gen-zero reports ✔ Connected · Type: sse. Keep --transport sse: without it the CLI registers a local stdio command and the server never connects.

Codex CLI, self-hosted Rust (gen-zero serve --mode sse) and the Python package: /agent-install.md and /llms-full.txt §5.

06 · Release Log

The MCP server reports gen-zero 0.1.0, and that is the only engine release so far. The entries below are dated changes to the API surface and this site, each with its commit. A new engine version will appear here when the server reports it.

  1. gen-zero 0.1.0Current · liveserver-reported version
    • Feature · Three MCP tools: zero (12 verbs), causal_fold, pipeline. Protocol 2024-11-05 over SSE.
    • Feature · pipeline decide modes: auto (K-MoE router), mcts, mpc_cem, astar, manifold_gflownet, cfr_nash, reflex. World models: residual, symplectic, contact.
    • Performance · Tier 0 Proceed gate path: P99 ≤ 80 µs (self-reported, /llms.txt). Scoring is prefill only: 0 generated tokens per decision.
    • Safety · Four-tier fail-closed gate (Proceed / Confirm / Escalate / HardStop). Non-finite, out-of-range or unknown input is refused with a typed code, never clamped.
    • Safety · Token auth on the public endpoint: a missing token gets HTTP 401. The sandbox cannot touch the caller's files or processes.
  2. Site · Wasm game engine2026-09-30167485d · 00e6eaf
    • Feature · Sokoban, Snake and 2048 demos on a Rust → Wasm engine (games-wasm 0.1.0, under 150 KB). They run in the browser and make no calls to the MCP server.
  3. Site · Quant risk-gate demo2026-09-30b1cb7e8 · b89d735
    • Feature · Jump-diffusion market simulation with a five-rule pre-trade gate, on the home page benchmarks section.
  4. API · Endpoint moved to api.gen-zero.ai2026-09-2993ec01b
    • Breaking · Public endpoint changed from https://gen-zero.ai/sse to https://api.gen-zero.ai/sse. Update saved client configs.
    • Feature · The old host still forwards /sse, /message, /health, /ready and /metrics to the new one, so same-origin browser probes (like §01) work.
  5. Research · Breakthroughs portfolio2026-09-29645d521
    • Performance · Folded adapter head reports 7.8–12.8 µs median per decision; CP-SAT safe dispatch reports 0 violations in 5,000 adversarial tests. Self-reported, not yet replicated.