GEN-ZERO Home Papers Breakthroughs Cases
[||]

CASE STUDY 01 · REAL-TIME NEUROSYMBOLIC CONTROL

The Tetris Proof: Real-Time Neurosymbolic Superiority

Why Combinatorial Real-Time Control Breaks Generative LLMs — And How Gen-Zero Answers It Today with Formal Safety Gating, with a Microsecond Reflex on the Roadmap.

Date: 2026-09
Latency (p50, live in this browser): not measured yet
Inaccessible holes at 40 lines (live): not measured yet
Token Burn: 0 Tokens ($0.00)

Interactive Arena

Autoplay shows the reference controller. Play yourself hands you the keys; the gate then only counts the moves it would have vetoed. Click the benchmark button below to run 20 fixed seeds (1000–1019) to 40 lines on your CPU and write the results into the meta bar and Section 3.

1. The Combinatorial Trap: Where Autoregressive Generation Hits Physics

A branching factor that compounds

With hard drops from spawn and no spins, a piece has at most 34 distinct placements on an empty 10-wide board (T, J and L), 17 for I, S and Z, and 9 for O. The mean over the seven pieces is about 23. These counts come from the arena engine itself and are pinned by a unit test. Looking four pieces ahead therefore spans up to 34⁴ = 1,336,336 placement sequences, before the random pieces that follow. Offline Tetris is NP-hard even to approximate [1], so any real-time controller must prune, and must prune well.

The physical clock

At 60 Hz a frame lasts 16.7 ms. Under 20G gravity the piece reaches the stack within a single frame, so the only time left is the lock delay: a fraction of a second in guideline games, less in arcade 20G modes. A controller that writes its move as text pays for that move token by token:

tdecision ≥ tfirst token + Nout / rdecode + tnetwork

This is arithmetic, not a measurement. Take a move written as about 40 output tokens and a decode rate of 50–100 tokens per second: decoding alone costs 0.4–0.8 s, before the first-token wait, the round trip and any hidden reasoning tokens, which reasoning models spend by the hundreds or thousands per step. Even under these favourable assumptions one decision spans dozens of frames. We did not time any hosted LLM for this page; replace the assumptions with your own numbers and the bound follows.

Spatial rules the decoder must re-derive every step

A text-generating agent predicts its next token from a serialized board. Nothing in that loop checks that the named placement is legal, reachable through the SRS kick table, or free of new holes: collision, rotation and line clears are exact discrete rules, and the model has to reconstruct them from text on every step. We have not measured an LLM's illegal-move or hole rate here, so we make no numeric claim about it. The engineering conclusion does not need one: whatever proposes a move, neural or symbolic, needs an exact checker downstream. The arena implements that checker.

2. The Dual-Track Neurosymbolic Architecture

Gen-Zero splits control into a fast neural proposer and an exact symbolic verifier, with a search between them that never leaves structured state. The badges say what exists today and what does not.

2. The Dual-Track Neurosymbolic Architecture BOARD STATE + NEXT QUEUE TRACK 1 · MAMBA-2 REFLEX target < 0.2 ms · not implemented ROADMAP TRACK 2 · LOOKAHEAD SEARCH beam 16 · depth 3 · zero tokens IN ARENA RANKED PROPOSALS TRACK 3 Δholes ≤ 0 IN ARENA veto → next proposal TRACK 3 · SAFETY GATE admit iff Δholes ≤ 0 pass PLACE PIECE no admissible move → least-bad move + breach counter (+ warning in autoplay)
View Mermaid Source
flowchart LR
    S["Board state + Next queue"] --> T1["Track 1: Mamba-2 reflex<br/>ROADMAP, not implemented"]
    S --> T2["Track 2: Lookahead search<br/>beam 16, depth 3, zero tokens"]
    T1 -.-> R["Ranked proposals"]
    T2 --> R
    R --> G{"Track 3: gate<br/>admit iff delta_holes <= 0"}
    G -- pass --> A["Place piece"]
    G -- "veto: next proposal" --> R
    G -- "no admissible move" --> B["Least-bad move<br/>+ breach counter (+ warning in autoplay)"]
    B --> A
    classDef roadmap fill:#12151A,stroke:#8E949A,color:#8E949A,stroke-dasharray:5 4;
    classDef live fill:#0e1218,stroke:#2FE3F0,color:#ECE7DC,stroke-width:2px;
    classDef gate fill:#0e1218,stroke:#FFA51F,color:#ECE7DC,stroke-width:2px;
    class T1 roadmap;
    class S,T2,R,A live;
    class G,B gate;
Track 1 · Design target · Not implemented

Mamba-2 sub-millisecond neural reflex

Intent: a selective state-space model [3] reads the board and scores surface bumpiness and open access channels for every candidate placement in under 0.2 ms. This is a target, not a measurement: no Mamba-2 model is trained or wired in. In the arena, the six-feature El-Tetris scorer (landing height, eroded cells, row and column transitions, holes, wells; weights from [2]) stands in for this track.

Track 2 · Implemented here as explicit-state search

Lookahead search with no token generation

The arena searches the current piece plus the next two previews (depth 3) with a beam of 16 boards per ply. Placements are those reachable by rotating at spawn, sliding and hard-dropping; tucks and spins are not searched. No token is produced at any point. The latent-space version of this search, over learned state embeddings, is roadmap.

Track 3 · Implemented here as an exact check

Formal safety gate: the zero-new-holes invariant

Let H(b) count the empty cells that have a filled cell above them in the same column. The gate admits a move a on board b only if H(place(b, a)) − H(b) ≤ 0, and walks down the search's ranking until a move passes. This is an exact integer check, not a solver. When no legal move passes, a piece still has to land: the arena plays the least-bad move and counts a breach; the count is shown on screen, and autoplay also logs a console warning. The invariant is enforced whenever it is feasible; it is not guaranteed, and the page reports every time it fails.

Gen-Zero's separate CP-SAT gate (Google OR-Tools [4]) belongs to its dispatch benchmark and is not wired into this arena. In the 2026-09 B3 evidence run it rejected 242 of 242 violating proposals across 1,000 solves, with a p99 solve time of 9.12 ms; 103 of those solves returned UNKNOWN and took a fallback, and the overall B3 acceptance verdict was NOT_ACCEPTED. Evidence bundle: Gen-Zero repository, benchmarks/results/b3_evidence.

3. Empirical Benchmark Scorecard

The Gen-Zero column fills in when the benchmark in the arena finishes. The other columns are empty on purpose.

MetricGPT-4o AgentClaude 3.5 Sonnet AgentDeepSeek-R1Gen-Zero reference controller (live, this browser)
Decision latency, median (p50)Not benchmarkedNot benchmarkedNot benchmarkednot measured yet
Decision latency, 99th percentile (p99)Not benchmarkedNot benchmarkedNot benchmarkednot measured yet
Decision latency, batch meanNot benchmarkedNot benchmarkedNot benchmarkednot measured yet
Decision throughput (pieces/s, compute only)Not benchmarkedNot benchmarkedNot benchmarkednot measured yet
40-line sprint completion rateNot benchmarkedNot benchmarkedNot benchmarkednot measured yet
Mean inaccessible holes at end of runNot benchmarkedNot benchmarkedNot benchmarkednot measured yet
Gate breaches (sum over 20 runs)Not benchmarkedNot benchmarkedNot benchmarkednot measured yet
Tokens and cost per gameNot benchmarkedNot benchmarkedNot benchmarked0 tokens · $0.00 (no network calls)

Why the LLM columns are empty. We have not run these agents under a common harness, so we publish no numbers for them. Filling the cells from vendor pages or intuition would be the kind of claim this page exists to refuse. A harness that drives each agent through the same 20 seeds and the same gate is future work.

Method. Seeds 1000–1019, seeded 7-bag, stop at 40 lines or 1,000 pieces. Search depth 3, beam 16, then the Δholes ≤ 0 gate. Decision time is performance.now() around search plus gate. Browsers coarsen that clock (the arena shows the resolution it measured), so p50 and p99 may be quantized; the batch mean (total time ÷ decisions) is not. Throughput here is compute-only: decisions per second of CPU time, not the animated play speed. Results depend on your device.

4. From Blocks to the Industrial World

What carries over is the pattern: a fast proposer, a search that stays in structured state, and an exact gate that can say no. Gen-Zero is not deployed in any of the domains below. Each card names the gate the domain would need and what this page does not show.

Finance

High-frequency trading

A proposer ranks orders; the gate checks pre-trade risk limits (position, notional, price bands) as exact constraints before an order leaves. Not shown here: the budget there is microseconds, and this controller runs in milliseconds of JavaScript; it would need a native port and exchange-grade testing.

Robotics

Agile robots and quadrupeds

A perception-to-motion reflex proposes the next action at control rate; the gate enforces joint limits and contact and stability constraints. Not shown here: hard real-time needs bounded worst-case execution time and certification. A browser p99 is not a worst-case bound.

Infrastructure

Datacenter power and traffic scheduling

A proposer suggests placements and throttles; the gate rejects any commit that creates a cycle in the wait-for graph or breaks a power cap. Gen-Zero Paper 2 studies deadlock avoidance in a synthetic environment. Not shown here: any production scheduler.

References

  1. [1] E. D. Demaine, S. Hohenberger, D. Liben-Nowell. Tetris is Hard, Even to Approximate. COCOON 2003.
  2. [2] I. El-Ashi. El-Tetris: An Improvement on Pierre Dellacherie's Algorithm. 2011.
  3. [3] T. Dao, A. Gu. Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality. ICML 2024.
  4. [4] L. Perron, F. Didier. CP-SAT. Google OR-Tools.
  5. [5] Tetris Guideline. Super Rotation System (SRS) wall-kick tables.

Cite this case study

@misc{genzero2026tetris,
  title        = {The Tetris Proof: Real-Time Neurosymbolic Superiority},
  author       = {{Gen-Zero Research Team}},
  year         = {2026},
  month        = sep,
  howpublished = {\url{https://gen-zero.ai/cases/tetris.html}},
  note         = {Case Study 01. Reference symbolic controller measured live in the
                  reader's browser; LLM baselines not benchmarked}
}