← Gen-Zero HomeTech Decoded中文版日本語PDF: Coming Soon
PAPER 1 / THE DECISION SERIES
An interactive research companion
View the source-audited Gen-Zero architecture and system flow

A decision.
Without a word.

Gen-Zero v0.1.0 moves from small-model probes to Qwen3.5-9B candidate readout, set attention and planning gates. Explore what changed—and what the early experiments taught us.

Explore the token tax
One input, two paths to a decisionThe hidden state can feed a decision head directly, or generate a succession of reasoning tokens before answering.Your questionHidden representationafter the input forward passDIRECT READOUTAUTOREGRESSIVEDecision headNext tokenrepeatYes / NoYes / No0 generated
Zero output tokens. Still a real computation.
The input and the decision head both have a cost.
The idea in one sentence

Read a decision from what the model already represents, instead of asking it to generate a written explanation first.

From a probe to a decision system.

The early 0.5B / 2B archives tested what a small readout could recover. The v0.1.0 implementation builds a richer candidate decision route: Qwen3.5-9B Vision Backbone + Permutation-Equivariant Set Attention + Planning MoE / CP-SAT Gating. Architecture has advanced; the archived scores remain historical measurements, not release benchmarks.

  1. Shared visual prefill

    Qwen3.5-9B encodes the image and prompt into a cached context. Input processing still costs time.

  2. Candidate direct readout

    Score supplied candidates from model logits without autoregressively generating an answer or scratchpad.

  3. Set attention

    The configured dual-head route provides candidate self-attention and state cross-attention. Equivariance concerns ordering; it does not guarantee semantic correctness.

  4. Planning & constraints

    The configured Planning MoE and CP-SAT gate control selection. Constraints, solver availability and route configuration define what is checked.

Implementation detail: visual logit scoring and set attention are separate calls. The visual scorer does not run DeepSetAttentionHead; decide_visual hands a state bundle to decide, whose configured dual-head route can use set attention. The code does not directly wire visual_hidden into that head. The map shows module responsibilities and routing, not a single tensor flowing through every box.

Zero generated tokens ≠ fixed total depth. Planning can add dependent computation internally. The paper’s conditional circuit bound applies only when the entire pipeline meets its fixed-depth assumptions. Set equivariance also does not remove label bias or all tie-breaking effects.

A result must say where it came from.

weights_loaded_from_checkpoint becomes true only after a strict dual-head checkpoint load. Read it alongside scorer, degraded and degraded_reason; the CPU scorer has a separate scorer_synced_from_checkpoint state. A fallback or random-initialized scorer must not be reported as a trained-model result. This flag describes the dual head, not authentication of every backbone weight or a retroactive repair of the old archives.

The vision engine separately exposes is_loaded / is_real and rejects scoring without real VLM weights. Inspect formal_verification on the underlying decision route to distinguish cp_sat_verified from an unavailable solver or a Python predicate filter. decide_visual does not forward every underlying verification field.

What rigor changed across the series

Paper 2 · failure found, isolated, closed. The Rust planner’s F01–F09 fail-closed regressions reject invalid transitions and rewards, preserve gate context, and recheck selected actions. The documented acceptance has 11 regression tests passing in both debug and release. The historical NaN fail-open finding is not the current Rust production behavior. These checks do not certify a physical actuator or every separate Python model entry point.

Paper 5 · measure the route you actually run. The supplied release brief gives roughly 1–3 ms as a native Rust production-path reference. The checked-in ci_logs/rust_test_verbose.log confirms 31 production-pipeline tests passed in 0.10 s; another recorded run took 0.15 s. No per-request benchmark was found to establish 1–3 ms, so it remains an unverified reference, not a measured latency guarantee. That suite duration is aggregate test evidence, not a per-request latency distribution. The old 746.7 ms Python/model measurement is not the current native-path latency; neither number measures the complete Qwen3.5-9B visual route. That route returns visual_prefill_ms, scoring_ms and latency_ms.

Release tag plus inspected local edits: gen-zero v0.1.0 @ dce82330da45; python/gen_zero/client.py (decide_visual, checkpoint loading and result metadata); python/gen_zero/model/dual_head.py; crates/gen-zero-planner/tests/fault_regression_tests.rs; docs/architecture/fault_regression_evidence/README.md. The widgets below are teaching simulations or archive visualizations; none runs the released model in your browser.

Contractive latent updating: a testable route

A proposed recurrent hidden-state update can add input-dependent depth without generating reasoning tokens. This is a formal design, not evidence that the current production decision route trains or executes this updater.

zₜ₊₁ = LayerNorm(zₜ + αₜ · Δₜ)Δₜ = σ(W₂ GELU(W₁ zₜ))Lip(U) ≤ γ < 1 ⟹ ‖zₜ₊₂ − zₜ₊₁‖ ≤ γ‖zₜ₊₁ − zₜ‖
Show derivation and evidence limits

Banach convergence is conditional: U must be a self-map on a complete metric space and its Lipschitz constant must satisfy γ < 1. LayerNorm and a learned gate do not automatically enforce this bound. Under that assumption, successive update distances decrease geometrically.

Theorem 3 · conditional expressivity

If the full update family uses R = n dependent steps and can implement the required length-dependent composition, its depth is no longer uniformly constant. The fixed-depth TC⁰ bound for a single-pass decoder then does not apply. This alone does not prove an NC¹ separation, an implemented model, or benchmark superiority.

CLUTRR · 10-hop relation composition

S0 · Left fold

26.23%73.77% refusal · unverified S0

S2 · Bidirectional meet-in-middle

64.75%35.25% refusal

S3 · CYK path lattice / chart fold

99.18%121/121 correct · 0 errors · 0.82% refusal

Historical Gen-2 reference, 122 test rows; accuracy includes refusals. S3 answered 121/121 correctly, made 0 errors, and refused 1 structurally ambiguous row (0.82%). The source README says its original rerun log is absent here. It reports 1.05 s for the test binary across hops, so a sub-second full-suite claim is not established.

MultiNLI · latent iteration

90.30% is a single-pass linear-probe comparison reported elsewhere on this site. No paired MultiNLI updater run or per-step residual trace was found in the checked evidence. γ ≈ 0.42, 5–8 steps and ‖Δₜ‖ < 10⁻⁴ are proposed targets, not measured convergence.

Analytic example only: normalized residual bound γⁱ with γ = 0.42; no measured MultiNLI trace.

Evidence boundary: research README (historical CLUTRR table, missing rerun log); site smoke-preview scorecard (90.30% probe). Neither is paired updater evidence.

Why write an essay
to say “yes”?

A binary answer carries at most one bit of information. Yet a reasoning model may generate hundreds of tokens before making that decision. Those tokens are a computational bill—and sometimes useful working space.

Autoregressive generation is a succession of steps: produce a token, feed it into the growing context, produce the next. A chain of thought can preserve intermediate results so that later steps can use them. Even with a cache, each new token needs another decoding step.

A zero-token readout takes a different route. Process the input, collect its hidden representation, and let a small decision head choose a label. If the needed distinction is already accessible, writing it out may be unnecessary. If it is not, a cheap readout does not magically supply the missing computation.

Write, then decide

Inputt₁t₂… tₙYesMore dependent computation; more decoding time

Models do not always need hundreds of tokens for binary decisions. This is a trade-off to test, not a universal requirement.

Read, then decide

InputHidden stateHead → YesNo generated scratchpad; no guarantee of correctness

“Zero-token” means zero generated output tokens. It does not mean zero input tokens, zero latency, or zero energy. The one-pass diagram is conceptual: the CPU candidate-selection implementation also runs candidate continuations from the prompt’s cached state before scoring them. Zero generated tokens does not imply one backbone call.

Feel the token tax.

Drag the assumptions. Compare one direct readout with 50 or 300 generated reasoning tokens. This is a transparent toy model of serving cost—not a benchmark from the paper.

Additional experiment: configurable dollar cost

Latency & cost playground

Illustrative · equal input
Shared prefill time for all three paths.
Constant rate; real rates vary with context and load.
Additional time on the zero-token path only.
Hypothetical dedicated worker rental, not API pricing.
0 tokens
0.205 s
50 tokens
1.200 s
300 tokens
6.200 s
One serial worker, no batching; cost per 1,000 decisions.
PathLatencyCompute costDecisions/min
0 tokens0.205 s$0.114292.7
50 tokens1.200 s$0.66750.0
300 tokens6.200 s$3.4449.7

See the equations and what is left out

Let P be input processing in seconds, H the head time, R the decoding rate, and N the number of reasoning tokens. T₀ = P + H; Tₙ = P + N/R. Serial throughput = 60/T decisions per minute. Cost per 1,000 decisions = 1,000 × T × hourly price / 3,600.

The comparison excludes final answer-token generation, network delay, queues, batching, tokenization, candidate preparation, and idle capacity. Equal input time and equal hardware price are assumptions. It estimates neither API bills nor energy. A faster path is only useful if its decisions are sufficiently accurate.

A factory with a fixed
number of stations.

Imagine a factory that may get wider as the input grows, but never adds more sequential assembly stations. That is the intuition behind constant depth.

Many workers. The same number of stages.

Within a stage, lots of workers operate in parallel. In TC⁰, threshold gates can inspect many incoming bits and ask questions such as “are at least half of these on?” Polynomial size limits how fast the number of workers can grow.

“Uniform” means there is a systematic recipe for wiring each factory size. It is not a separate factory full of secretly precomputed answers for every input.

Depth is the chain of dependencies.

A fixed stack of transformer layers suggests this assembly-line picture. Generating another token reuses that stack, letting a later computation depend on a newly written intermediate result.

A transformer layer is not literally one threshold gate. The paper assumes a constant-depth threshold simulation of the whole pipeline, with bounded precision and polynomially bounded encodings. A model’s name alone does not establish that assumption.

Additional experiment: single-pass layer trace

Trace a signal through the layers

Conceptual diagram
Transformer layer information flowFour illustrative layers feed a decision head. Use the controls below to trace one pass or repeated scratchpad passes.
Stage 0 / 5

The input is ready. Advance to watch information move through four illustrative layers and a decision head.

The four layers are schematic, not the architecture of either evaluated model. In the scratchpad view, each new token is added to context; cached states can be reused. Repetition adds dependent computation, not an automatic accuracy guarantee.

Parity: a coin-flip problem that fits.

Start with a coin showing tails. Each 1 flips it; each 0 leaves it alone. Is it heads at the end? That depends only on whether the number of 1s is odd.

You could flip the coin sequentially. But threshold gates can test all possible count thresholds in parallel, identify an odd total, and combine the results in a constant number of stages. Parity is in uniform TC⁰. State dependence alone does not make a task hard.

Try it: switch the input bits

3 flips · odd · heads

This tiny example illustrates the rule. The paper’s construction applies to arbitrary input length, with a fixed depth and polynomial size.

Where the easy shortcut runs out

Now imagine a sequence of instructions that reshuffles five objects. To determine the final arrangement, the order of the instructions matters. The paper studies a specific permutation-product task with 120 possible states—not every state machine.

The limit is conditional: assuming the nonuniform classes TC⁰ and NC¹ differ, a fixed-depth threshold-simulable pipeline cannot solve this specified task correctly at every input length.

A recurrent machine can update the state instruction by instruction. A parallel tree can also combine transitions using depth that grows logarithmically with sequence length. Written chain of thought is one way to add computation; it is not the only way.

What this theorem does—and does not—say

The separation TC⁰ ≠ NC¹ is an assumption, not a proved result here. The conclusion concerns an asymptotic family of problems and specified computational assumptions. It does not prove that a finite benchmark is impossible, that all state-dependent tasks are hard, or that a trained model will learn the constructive recurrent algorithm. It also does not explain an individual arithmetic error.

Read a signal, without writing a token.

This historical teaching probe makes a simple readout visible; it is not the v0.1.0 set-attention architecture. These synthetic activations are a teaching model, not measured neural activations or benchmark results.

Amber: negativeCyan: positive

Each column is a feature. Each row is a schematic layer. The probe reads only the final row: score = mean(h). The sign selects a label; it is not a probability.

Synthetic residual activation heatmapFour schematic layers of eight synthetic features. Move the slider to change their signed values and the final readout.Readout

Early prototypes:
the prediction ledger.

These early 0.5B / 2B prototype archives establish empirical boundaries, not v0.1.0 performance. Their different sample mixes and unresolved historical execution provenance remain part of the record. The new candidate architecture is described above; it has not been rerun on these examples for this exhibit.

Qwen3.5-2B · GPU archive

65.13%

Macro accuracy across 13 tasks, 390 examples. Each task contributes 30 examples.

254/390 correct; macro and micro accuracy are both 65.1282% (65.13% rounded). Prediction rows and error counts reconcile. Full prompts and authenticated execution-to-checkpoint linkage remain unavailable. This is an archive claim, not certified model performance.

Qwen2.5-0.5B · INT8 CPU archive

32.55%

Macro accuracy across 13 tasks, 930 examples. A trained low-dimensional head reads the frozen trunk.

357/930 correct; micro accuracy is 38.3871% (38.39% rounded), while macro accuracy is 32.5513% (32.55% rounded). The recovered ledger matches the counts, but inference was not replayed. Macro uniform chance is 33.16%; the macro majority baseline is 48.08%.

Macro ≠ micro. Macro gives every task equal weight. Micro gives every example equal weight. The 0.5B archive has 400 PAWS and 200 GSM8K examples, so those tasks dominate its micro score. Neither headline establishes a scaling law or a matched advantage over chain of thought.

The “yes-man”
failure mode.

PAWS asks whether two sentences mean the same thing. Similar words can hide different meanings: “The dog chased the cat” and “The cat chased the dog” share their vocabulary, but reverse who did what.

Sentence pair above is an invented illustration, not an archived test row.

In the 0.5B CPU ledger, the system calls all 400 pairs paraphrases—a 100% single-class prediction collapse. Half actually are. Half are not. A 50% accuracy score conceals a classifier that never recognizes a negative.

400 answers.
One repeated label.

This historical shortcut-learning warning is observed prediction collapse, not proof of identical hidden representations or a known internal cause. Projection, head training, candidate preference, quantization, and prompts remain competing explanations.

Same examples. Two very different pictures.

ParaphraseNot a paraphrase

Each square represents 4 examples. Grouped for clarity, not in original row order.

Actual labels: 200 paraphrases, 200 non-paraphrases.

Confusion matrix · 0.5B CPU PAWS slice
Actual classSays yesSays no
Paraphrase2000
Not paraphrase2000

Non-paraphrase recall: 0%. Every negative example is missed.

A scratchpad has a job.

Multi-step math asks a model to preserve intermediate quantities and use them correctly. Removing the written working space can remove a useful route to the answer.

Imagine a stall with 18 eggs. It keeps 3, uses 5 for baking, and sells the rest for $2 each. The working is compact, but each step carries information into the next.

18 − 3 = 15Account for the eggs kept.
15 − 5 = 10Carry that result into the next subtraction.
10 × $2 = $20Turn the remaining stock into revenue.

Illustrative problem; not a scored item or a reconstruction of the archived prompt.

Early GSM8K: an empirical boundary

Four-choice candidate selection, not free-response GSM8K.
ArchiveCorrectAccuracy
2B GPU3 / 3010.00%
0.5B CPU53 / 20026.50%

Uniform four-choice guessing has 25% expected accuracy. The CPU’s 95% Wilson interval, 20.87–33.02%, includes that baseline.

A failure is not an impossibility proof.

These configurations perform poorly. But the archives do not compare the same model with and without a scratchpad. They cannot establish that missing chain of thought caused the errors, or that every forward-pass model must fail at arithmetic.

The GPU slice puts every gold answer in position A, so always choosing A would score 100%. The CPU slice balances A–D. A GPU error record also has an incomplete prompt preview of unresolved origin. Different label placement, sample sizes, and protocols prevent a paired comparison.

Less writing.
Still more
to prove.

A hidden-state decision head can be useful when the representation already carries an accessible answer. That is a narrower—and more testable—claim than “reasoning for free.”

The decisive next experiment would hold the examples, prompts, candidates, and model fixed, then compare direct readouts, candidate scoring, generated reasoning, and trained latent recurrence. Measure accuracy alongside all backbone calls, generated tokens, and end-to-end latency.

The question is not just “How fast did it decide?” It is “What computation made that decision reliable?”

Source snapshot: v0.1.0 / HEAD = dce82330da45. The inspected checkout also contains local Python edits, including finer scorer-provenance reporting; those additions are not all part of the immutable release tag. No model inference or Rust benchmark was rerun for this exhibit.

Source notes & reading guide

Paper 1 — Zero-Token Decision Circuits: Expressivity Bounds and Empirical Limits of Forward-Pass Gating in Autoregressive Models.

This companion follows the supplied draft.md and paper.tex. It adds explanatory analogies and interactive illustrations, not new model experiments.

All visual assets, styles, fonts, and scripts are embedded. Open this file directly in a browser; no connection is needed.

Theory: §3, Proposition 3 (parity), Proposition 4 (pipeline closure), Theorems 1–2 (conditional tracking limit and recurrence).

Scores: §§5.2–5.3; macro versus micro in §5.6. Failures: §§6.1–6.2. Provenance: §6.5 and the appendix claim-to-evidence map.

Audit artifacts: evidence/audit.py, evidence/revision_checks.py, and evidence/failure-mode-diagnostics.json. Count reconciliation is not a rerun of inference.

Gen-Zero model architecture
and neural pipeline

The v0.1.0 vision route uses Qwen3.5-9B candidate scoring and configured downstream decision modules. Historical alignment and semantic-risk routes remain separate. The spatial world model is a separate learned branch. Explore the layers and inspect their tensor shapes.

Code audit: · gen-zero @ dce82330da45. Solid arrows show implemented data flow; dashed arrows show a proposed dispatch contract. These components are not one universally deployed serial pipeline.

Implementation detail: visual logit scoring and set attention are separate calls. The visual scorer does not run DeepSetAttentionHead; decide_visual hands a state bundle to decide, whose configured dual-head route can use set attention. The code does not directly wire visual_hidden into that head. The map shows module responsibilities and routing, not a single tensor flowing through every box.

All flows visible. Select a layer to inspect its dimensions and evidence.

On a narrow screen, scroll the diagram horizontally. Every layer is keyboard selectable; the source notes also provide a text version.

Gen-Zero multi-tier architecture with inspectable dimensions The vision route uses shared prefill and candidate scoring with configured downstream planning. Separate archived backbone features feed alignment and semantic-risk branches. Independent 64-dimensional spatial features and 16-dimensional actions feed residual dynamics. Rust transition checks are implemented and regression-tested; dashed arrows refer only to separate physical dispatch integration. Select a node for details. v0.1.0 vision backbone + research archivesVision backbone + research archives v0.1.0: Qwen3.5-9B vision · archives: 0.5B / 2B / 72B / 70B Visual and archived research routes are distinct; inspect each node Spatial inputs · d=64 + 16Spatial observations Engineered state z: [B, 64] Action vector a: [B, 16] Candidate direct readout · cached visual contextVisual scoring / archived readout Vision: cached context → candidate logits Archives: pooled hidden-state features Semantic scoring · scalar riskFast semantic risk gate Frozen 0.5B · 16-shot ICL 3 orders · differential PMI Residual spatial dynamics · 80 → 64Residual world dynamics [B, 80] → Δz: [B, 64] z′ = z + f(z, a) Permutation-equivariant candidate attentionSet attention + gating Candidate self- / cross-attention Planning MoE / CP-SAT Alignment and fusion are separate routesCross-model alignment Paired SVD / Procrustes Supervised heads + logit fusion Live thresholds and diagnostic τ=0.50Semantic policy tiers Escalate ≥ 0.4494 Hard stop ≥ 0.7620 F01–F09 · Rust fail-closed regressionsFail-closed boundary F01–F09: Rust checks repaired Reject invalid model transitions Result provenance · inspect the actual routeResult + provenance scorer · checkpoint state degraded · degraded_reason Evaluation · selected configurationsSelected test outcomes PubMedQA: 78.40% · 196/250 Aegis: 81.60% · 204/250 Risk evidence · small curated suiteStored risk evaluation AUC 0.944 · 36 requests τ=0.50: 18/18 danger recall Checked decisions · physical dispatch remains separateChecked dispatch contract Reject NaN / Inf outputs Latch halt until controlled reset Zero generated answer tokens still require input processing, scoring and real computation.

Inspect a model layer

Select any box to see its tensor dimensions, mechanism and implementation limits. Use Tab, then Enter or Space, or select with a pointer.

Implementation evidence and architecture boundaries

Backbone provenance. The visual entry point in client.py configures Qwen/Qwen3.5-9B and shares image-plus-prompt prefill. The 0.5B / 2B probe scores and 72B / 70B alignment features belong to separate research archives. Their dimensions and scores are not specifications or evaluations of the 9B release route.

Training scope. The Transformer weights remain frozen in these paths. The semantic gate uses 16 labeled in-context examples and differential log-likelihood (PMI), without fitting a safety head. The spatial model still has a learned outcome/reward head; supervised alignment heads also remain. “No backbone fine-tuning” does not mean no downstream training.

Safety scope. The Rust planner’s F01–F09 fail-closed regressions reject invalid transitions and rewards, preserve gate context, and recheck selected actions. The documented acceptance has 11 regression tests passing in both debug and release. The historical NaN fail-open finding is not the current Rust production behavior. These checks do not certify a physical actuator or every separate Python model entry point.

  • crates/gen-zero-model/src/choice_head.rs: canonical ActionId sorting and Simplex ETF geometry, Sₖ ⊂ ℝᵏ⁻¹. The retained direct Rust benchmark reports 0/16,232 identity flips (0.00%), with fixed representation and candidate identities; maximum probability drift is 5.96 × 10⁻⁸. Paper evidence: equivariant-choice-head/evidence/revision/results.json.
  • python/gen_zero/world_model/neural_dynamics.py: [B,64] state + [B,16] action → [B,80] concatenation → Δz=f(z,a) → z′=z+Δz; learned sigmoid outcome head. Dimensions come from the spatial configuration, not universal class constants.
  • benchmarks/suites/geometric_latent_fusion.py: orthogonal Procrustes diagnostic and paired-SVD core/residual features. benchmarks/suites/evaluate_manifold_pareto_ensemble.py: supervised normalized-logit fusion, distinct from geometric projection.
  • docs/zero/29-pubmedqa-aegis-unified-manifold-evaluation-closure.md: PubMedQA 78.40% (196/250), selected dual-head fusion; Aegis Track A 81.60% (204/250), selected single LLaMA head. These descriptive results do not establish statistical superiority.
  • python/gen_zero/service/semantic_risk.py and python/gen_zero/service/risk_data/report.json: frozen 0.5B, 16-shot ICL, three orders, overlapping windows, PMI scoring; AUC 0.944. Diagnostic τ=0.50 flags 18/18 dangerous and 7/18 benign requests. Live thresholds are 0.4494 / 0.7620.
Text version of every layer

v0.1.0 vision backbone + research archives

The visual entry point in client.py configures Qwen/Qwen3.5-9B and shares image-plus-prompt prefill. The 0.5B / 2B probe scores and 72B / 70B alignment features belong to separate research archives. Their dimensions and scores are not specifications or evaluations of the 9B release route.

Spatial inputs · d=64 + 16

These are engineered spatial features, not a projection from the Transformer hidden state. The evaluated spatial configuration concatenates z and a into [B, 80]. The class supports configurable state and action dimensions.

Candidate direct readout · cached visual context

decide_visual calls prefill_visual_context, then score_candidates_direct, then decide for configured planning and constraint gating. Supplied candidates are scored without autoregressive answer generation. The result reports direct_probs, visual_prefill_ms, scoring_ms and latency_ms; zero generated tokens does not mean zero inference cost.

Semantic scoring · scalar risk

For each request window, subtract empty-request log-odds from log P(" dangerous") − log P(" safe"). Average over three demonstration orders; use the riskiest overlapping window and apply sigmoid. No newly trained neural safety classifier is used for this semantic gate. The label logits come from the frozen model; it does not emit an explanation.

Residual spatial dynamics · 80 → 64

neural_dynamics.py uses a residual MLP with LayerNorm and GELU. A separate learned sigmoid outcome/reward head returns [B]; it remains in the current code and is distinct from the frozen semantic risk gate. The class default hidden width is 128 with two residual blocks; checkpoint configuration is authoritative.

Permutation-equivariant candidate attention

The v0.1.0 dual-head model contains DeepSetAttentionHead: candidate self-attention and cross-attention to state context. Candidate order equivariance does not prove semantic accuracy or eliminate label and tie biases. The Rust canonical ActionId / simplex ETF head is a separate mechanism; its shuffle benchmark is not a benchmark of the vision set-attention route. Implementation detail: visual logit scoring and set attention are separate calls. The visual scorer does not run DeepSetAttentionHead; decide_visual hands a state bundle to decide, whose configured dual-head route can use set attention. The code does not directly wire visual_hidden into that head. The map shows module responsibilities and routing, not a single tensor flowing through every box.

Alignment and fusion are separate routes

Orthogonal Procrustes uses R=UVᵀ from the SVD of XᵀY. GeometricLatentFusion stores that map as a diagnostic; its features use paired SVD axes, scale-matched core averages and residuals. Supervised heads learn from labels. The winning PubMedQA path combines normalized head logits, not Procrustes coordinates: 0.75 Qwen + 0.25 LLaMA.

Live thresholds and diagnostic τ=0.50

The checked-in service uses two boundaries: 0.4494 for escalation and 0.7620 for a hard stop. τ=0.50 is the paper’s binary analysis threshold, not the live policy. On the 36 stored cases it flags 18/18 dangerous and 7/18 benign requests; the live tiers yield 12 stops + 6 escalations for dangerous cases and 9 escalations for benign cases.

F01–F09 · Rust fail-closed regressions

The historical Rust NaN fail-open findings were isolated and fixed. F02/F03 reject non-finite transition rewards and successor coordinates; F06 checks hazards during selection, search and final prediction; F09 rejects invalid entropy. Eleven regression tests cover F01–F09 in debug and release. This is a software boundary, not a physical motor-stop certification.

Result provenance · inspect the actual route

Read weights_loaded_from_checkpoint together with scorer, degraded and degraded_reason. The vision engine separately requires real loaded weights; is_loaded / is_real report its state. The dual-head checkpoint flag alone does not prove that a trained neural scorer produced this particular result. The separate historical Rust ETF benchmark recorded 0/16,232 identity flips; that result does not measure visual candidate accuracy.

Evaluation · selected configurations

PubMedQA selected fuse0.75+bbp|raw (dual-head normalized logit fusion). Aegis Track A selected llama+bbp|raw (single-model head). These are descriptive local test results, not verified SOTA or a universal alignment gain. The two-task macro bootstrap interval includes zero.

Risk evidence · small curated suite

AUC is 0.944444 from 18 dangerous and 18 benign stored requests. At τ=0.50 dangerous recall is 18/18, with seven benign flags. The bare chmod regression probe remains a miss. The implementation records real computation time; cached prompts reduce repeated setup but do not remove inference latency.

Checked decisions · physical dispatch remains separate

The repaired Rust production path validates transitions and policy before returning a decision. A persistent hardware dispatch latch is a separate integration contract. Do not conflate the closed Rust planner defects with the independent Python neural_dynamics.step method or claim that a decision ledger proves an actuator executed safely.

Gen-Zero system architecture & flow

The Rust runtime connects entry points, policy, planning, decision heads and cryptographic audit. The map below groups their responsibilities; the selected verb determines the actual call path.

Source audit: 29 September 2026 · gen-zero@dce82330da450d9938822f9db37c61cd29be2e49. Source-verified wiring; no new deployment or model evaluation is claimed.

Five layers of the Gen-Zero Rust runtimeA responsibility map, not a mandatory five-stage pipeline. Hover, focus or click a layer for source-backed details. Full text is also available below.Entry layergen-zero-cli · gen-zero-serviceCLI · JSON-RPC · MCPGate & securitygen-zero-gateProceed · confirm · escalate · hard stopPlanning & dynamicsgen-zero-planner · gen-zero-worldmodelNumeric rollouts and candidate filteringDecision coregen-zero-modelCanonical ActionId · Simplex ETFVerification & auditgen-zero-provenanceKeyed BLAKE3 · linked entries · MMR proofs
Hover or focus to preview. Click, Enter or Space pins a tooltip; Escape dismisses it. On touch screens, tap a layer.

Select a layer to inspect its implementation and limits.

What is not implemented here

The requested AST parser + Linux capabilities stack, physical collision checking, and persistent fail-closed latch are not an integrated execution path in these Rust crates. Invalid-input rejection and policy escalation exist; they do not establish a sandbox or a physical stop.

These service routes return decisions and simulation results. A decision ledger is not evidence of external action execution.

From request to response

  1. Enter and bind. The CLI or MCP/HTTP service receives a request. The router captures its mount snapshot, validates route-specific fields and selects the verb.
  2. Assess the request. Text ask, route and imagine obtain semantic risk from the Python bridge. Missing or invalid risk escalates; a hard-stop tier blocks that route. Policy constraints also apply when selecting actions.
  3. Use the requested route. Text ask uses semantic scoring, with an explicitly labeled first-feasible fallback when unavailable; unassessed risk still escalates. Text route reports unavailable when its bridge cannot run. Numeric planning uses gen-zero-planner and gen-zero-worldmodel; simulate, what_if and shadow audit expose untrained priors. An explicit ETF head uses canonical action IDs and simplex projection. Numeric cognitive requests have their own geometry verification path.
  4. Gate, record, return. Action-selection routes combine their policy verdict with request risk. Successful ask outcomes and pipeline decide results with a decision append to the provenance ledger before returning their audited outcomes. The response includes route metadata; it does not dispatch an external command.
Source map and full layer descriptions

Entry layer

gen-zero-cli · gen-zero-service — The CLI starts McpServer or submits a decision. The service accepts MCP over stdio and HTTP/SSE, plus REST routes. Its zero router binds an immutable mount snapshot and dispatches by verb; these are different routes through shared crates.

Under /ebs/pj/gen-zero/crates/: gen-zero-cli/src/main.rs; gen-zero-service/src/server.rs; gen-zero-service/src/zero.rs

Gate & security

gen-zero-gate — PolicyGate supports formal constraints, confirmation registration, entropy and semantic risk tiers. Its default has no registered constraints or confirmation actions. Text ask, route and imagine use the Python semantic-risk bridge; missing or malformed risk escalates. Shell AST parsing and Linux process capability enforcement are not wired into this Rust gate.

Under /ebs/pj/gen-zero/crates/: gen-zero-gate/src/policy.rs; gen-zero-gate/src/risk.rs; gen-zero-service/src/bridge.rs

Planning & dynamics

gen-zero-planner · gen-zero-worldmodel — Numeric latent requests can use MCTS, MPC-CEM or A*. The service exposes untrained residual or symplectic dynamics, with finite-input checks and terminal-state hazard handling. Planning filters gate-hard-stopped candidates; fixed-plan simulation still steps blocked actions and records their gate tiers. This is offline dynamics filtering, not verified physical collision checking. No persistent fail-closed execution latch is wired into this Rust route.

Under /ebs/pj/gen-zero/crates/: gen-zero-service/src/worldsim.rs; gen-zero-planner/src/pipeline.rs; gen-zero-worldmodel/src/dynamics.rs

Decision core

gen-zero-model — ActionETFChoiceHead sorts stable ActionIds, projects a shared representation onto a regular simplex, scatters scores to caller order and applies softmax. The core has a deterministic near-tie rule. The service exposes ETF through an explicit head option; ordinary text decisions use the semantic bridge, so ETF is not every request’s default head.

Under /ebs/pj/gen-zero/crates/: gen-zero-model/src/choice_head.rs; gen-zero-service/src/zero.rs

Verification & audit

gen-zero-provenance — DecisionAuditEntry includes the previous MMR root. A keyed BLAKE3 Merkle Mountain Range commits the decision history and supports inclusion proofs. Successful ask outcomes and pipeline decide results with a decision append records. Persistence via GENZERO_MMR_PERSIST_PATH is optional; the last 4,096 leaves retain inclusion proofs, and a trusted root must be retained externally; this is not an execution audit proving that an external shell command or motor action ran.

Under /ebs/pj/gen-zero/crates/: gen-zero-provenance/src/entry.rs; gen-zero-provenance/src/mmr.rs; gen-zero-service/src/zero.rs

Related runtime infrastructure includes gen-zero-core (shared types), gen-zero-lod (graph facts), gen-zero-storage (snapshots) and optional gen-zero-nanocore routes. The paper experiments below or above are research evidence, not a substitute for this runtime map.