What the public Gen-Zero MCP server runs right now, checked live from your browser. The real input schemas of its three tools, responses recorded from the live server, install commands we ran ourselves, and a release log with dates. Every number on this page names its source.
01 · Live Endpoint Status
This panel opens a real MCP session from your browser: GET /health, an SSE handshake on /sse, then initialize and tools/list. Nothing here is hard-coded. If a step fails, the panel says which step and why.
games-wasm 0.1.0 (Rust → Wasm, runs locally, not part of the MCP server)
Experimental preview. The public endpoint and the gz_public_free token are for research, benchmarking and evaluation. The endpoint is a read-only simulation sandbox: it cannot touch your files, processes or storage. No warranty, no commercial SLA.
02 · Measured Figures, With Sources
We list a latency only when a run or a file backs it. Each row says where the number comes from and how far it can be trusted. Different rows come from different runs, so do not add them up.
zero · ask with free-text state, scored by the native Qwen backbone plus risk classifier, one call on 2026-10-01
2 278 ms
Live response, _meta.scorer.forward_ms (sample in §03)
One sample
Generated tokens in the core decision path
0
Architecture: scoring is prefill only, no decoding. /llms-full.txt §6
Design property
Read before quoting. "0 generated tokens" is not "0 tokens processed": the semantic ask path still reads a prompt (20 prompt tokens in the sample below). No end-to-end speed or cost comparison against an LLM exists yet. "Zero hallucination" has not been measured, so we do not claim it.
03 · MCP Tool Reference
Schemas below are copied from the server's own tools/list answer. Responses are real calls to the public endpoint on 2026-10-01, trimmed only where a field holds 1024 floats. Call a tool with tools/call and {"name": ..., "arguments": {...}}.
zero
Polymorphic decision primitive
One entry point for 12 verbs. Name the verb in action, or leave it out and the server picks one from the request. Scoring is prefill only: the server reads your text and candidates but writes no tokens. Every refusal comes back as a typed error, never as a quiet default.
The ranker chose the safe option with 0.984 confidence. The gate still did not release it, because the request itself contains rm -rf /. A held action comes back as isError: true with the reason, so a client cannot mistake it for approval.
causal_fold
Multi-hop relation fold
Folds a chain of relation ids into one relation, using a composition table T(r1, r2, gender) = result that you supply. The default chart strategy tries every bracketing and answers only when exactly one relation survives at the root. With no table, or two surviving answers, it refuses. It never guesses.
Parameter
Type
Meaning
edges
int[1..64]
The chain: one relation id (0–65535) per edge.
genders
enum[]
Node genders along the chain, Male / Female / Unknown. Must have one more entry than edges.
axioms
object[≤4096]
The table: {r1, r2, gender, result: [ids]}. Missing means empty, so any chain of 2+ edges refuses.
{
"isError": true,
"content": [{ "type": "text", "text":
"CausalFoldRefused: no bracketing of the chain closes under the learned table" }],
"_meta": {
"reject": {
"code": "CausalFoldRefused",
"stage": "causal_fold",
"http_status": 422
},
"causal_fold": { "strategy": "chart", "edges": 2,
"axioms": 0, "step_failed": 2 }
}
}
Observed in our calls: an axiom's gender is matched against the node at the far end of the pair it composes. With genders [Male, Female, Male], the axiom for (1, 2) must say Male.
pipeline
Latent world-model pipeline
The Rust ProductionPipeline on a 1024-dimension latent state. Four operations: simulate a fixed plan, what_if over candidate first moves with trap detection, audit_action with a verdict of Approved / WarnHazard / RejectLethal, and decide. PolicyGate hard stops are pruned before any engine runs.
Parameter
Type
Meaning
oprequired
enum
simulate · what_if · audit_action · decide
staterequired
number[1024]
Latent state, exactly 1024 finite numbers. A divergent state is refused.
candidates
int[]
Candidate action ids for what_if and decide (at most 16 for decide).
action · actions
int · int[]
The action to audit, or the fixed plan to simulate.
horizon
int 1–128
Rollout length. Default 5 for what_if, audit_action, decide.
mode
enum
For decide: auto (K-MoE router), mcts, mpc_cem, astar, manifold_gflownet, cfr_nash, reflex.
entropy · warn_risk
0–1 · (0, 1]
Exploration level for decide; risk score above which audit_action warns (default 0.3).
{
"isError": false,
"content": [{ "type": "text",
"text": "audit_action: Approved (risk_score 0.0511)" }],
"_meta": { "engine": "production_pipeline", "pipeline": {
"op": "audit_action",
"audit": {
"verdict": "Approved",
"gate_tier": "Tier0Proceed",
"risk_score": 0.05105,
"warn_risk": 0.3,
"survival_horizon": 5,
"first_hazard_step": null,
"safety_calibrated": false,
"reasons": [
"no hazard in 5 step(s) of horizon 5; risk_score 0.0511 < warn_risk 0.3",
"safety estimate is uncalibrated (latent_norm_boundary_margin): a model-internal margin, not an observed failure rate"
]
} } }
}
The server states its own limit: the default world model is an untrained latent prior, and safe_prob is an uncalibrated margin, not a failure rate. Treat pipeline output as a structural check, not a forecast.
04 · Fail-Closed Policy Gate
Every decision passes a four-tier gate before it is released. Unknown, non-finite or out-of-range input does not fall through to a default: it is refused with a typed code.
Tier 0 · Proceed
Low risk, high confidence. The action is released.
Tier 1 · Confirm
High-impact or irreversible. Needs credentials or a second confirmation.
Tier 2 · Escalate
Uncertain. Goes to deeper lookahead or a human. Risk classifier threshold: p ≥ 0.4494.
Tier 3 · HardStop
Hard constraint broken or formally infeasible. Blocked. Classifier threshold: p ≥ 0.762.
Tier definitions: /llms.txt. Thresholds: Paper 5, confirmed in the live _meta.risk.thresholds field above. Known gap, stated in Paper 5: a bare chmod -R 777 / scores below the escalate threshold, so this classifier is one layer of defence, not the only one.
05 · Client Install & Config
Pick your client. Each block has a copy button. Restart the client after you edit a config file.
Verified 2026-10-01:claude mcp get gen-zero reports ✔ Connected · Type: sse. Keep --transport sse: without it the CLI registers a local stdio command and the server never connects.
.cursor/mcp.json in the project, or ~/.cursor/mcp.json for all projects
The MCP server reports gen-zero 0.1.0, and that is the only engine release so far. The entries below are dated changes to the API surface and this site, each with its commit. A new engine version will appear here when the server reports it.
gen-zero 0.1.0Current · liveserver-reported version
Feature · Three MCP tools: zero (12 verbs), causal_fold, pipeline. Protocol 2024-11-05 over SSE.
Feature · pipeline decide modes: auto (K-MoE router), mcts, mpc_cem, astar, manifold_gflownet, cfr_nash, reflex. World models: residual, symplectic, contact.
Safety · Four-tier fail-closed gate (Proceed / Confirm / Escalate / HardStop). Non-finite, out-of-range or unknown input is refused with a typed code, never clamped.
Safety · Token auth on the public endpoint: a missing token gets HTTP 401. The sandbox cannot touch the caller's files or processes.
Site · Wasm game engine2026-09-30167485d · 00e6eaf
Feature · Sokoban, Snake and 2048 demos on a Rust → Wasm engine (games-wasm 0.1.0, under 150 KB). They run in the browser and make no calls to the MCP server.
Site · Quant risk-gate demo2026-09-30b1cb7e8 · b89d735
Feature · Jump-diffusion market simulation with a five-rule pre-trade gate, on the home page benchmarks section.
API · Endpoint moved to api.gen-zero.ai2026-09-2993ec01b
Breaking · Public endpoint changed from https://gen-zero.ai/sse to https://api.gen-zero.ai/sse. Update saved client configs.
Feature · The old host still forwards /sse, /message, /health, /ready and /metrics to the new one, so same-origin browser probes (like §01) work.
Research · Breakthroughs portfolio2026-09-29645d521
Performance · Folded adapter head reports 7.8–12.8 µs median per decision; CP-SAT safe dispatch reports 0 violations in 5,000 adversarial tests. Self-reported, not yet replicated.