recorded replay · JEV playing an untouched holdout seed recorded replay · the standard game recorded replay · game variation 1: the shields start to move
  1. Can a fast model play Space Invaders? Yes.
  2. Can experience make it better? Yes.
  3. Can what it learned become wrong? Yes.
  4. Can it know when to let go?

That's where things got interesting.

Memory helps.Until the world changes.

No shared memory –
Mitosis memory –

stable world · average points below the best setup · lower is better

The memory didn't disappear.
The world it described did.

NOOP
recorded replay · every move chosen by JEV · re-derived exactly from the seed and action log · points

First, we had to survive one frame.

Space Invaders made latency visible.

JEV
ms / decision
Haiku 4.5
ms / decision

one second of play, slowed down · median per decision, holdout games

× faster

A language model could reason about the frame.
JEV could decide while the frame still mattered.

JEVper game
Haiku 4.5per game

average score · untouched seeds, run once

real-time head-to-head vs Haiku 4.5
turn-based vs Haiku 4.5 (5 seeds, cost-limited)
show receipts

Then the benchmark started arguing back.

one real frame · what the model was shown

RAW

        
Fire jitter fired with a shot already in flight reversed direction on of moves scored
COMPRESSED

        
Persistent motion fired with a shot already in flight reversed direction on of moves scored

Same model. Same game.
Different representation.
Different policy.

So we tried something embarrassing:
a dumb rule.

sweep bot (no model) vs JEV on the compressed view · 50 fresh seeds · points without the bonus ship

High score was not enough evidence of understanding.

every strong result got a cheaper control

    What survived: tell JEV which moves are safe, and to sweep unless it has a reason not to. On the holdout that player scored against the sweep bot's , and won seeds.

    Play

    What should I do now?

    ms per move

    Eval

    Did that decision survive reality?

    on unseen life-or-death moments

    Lab

    What deserves another experiment?

    which setup gets the next games, and which old finding to doubt

    One decision model.
    Three timescales.

    Let experience outlive the agent.

    Four workers search ways to play: stand still, track, or sweep · dodge or not · fire when ready, or aimed. Each round, each worker plays one real game.

    rounds later, we killed the workers.

    The new workers had never played those games.

    But the swarm had.

    MITOSIS · shared findings
    
    
        

    Tenki gave the swarm bodies.

    Mitosis gave it inheritance.

    This picture is the distributed run: real Tenki sandboxes. workers created, replaced exactly on schedule, unplanned deaths. Every newborn worker was checked at birth: it knew nothing. Findings shown are real inherited findings from trial 502.

    stable world · main run · paired trials · pre-registered

    No memory
    Shared memory

    average points below the best setup · lower is better

    Shared memory cut search error by more than half.

    /trials better with memory
    fewer games per trial spent re-testing losing setups the swarm had already tried
    1. a new worker picks a strategy
    2. checks the shared inheritance
    3. skips the experiment a dead worker already ran

    The swarm stopped paying twice for knowledge it had already earned.

    standard gamevariation 1 · the shields move

    Then we changed the world.

    old world · true mean score of each setup

      Some inheritance survived. Some became wrong.

      The difficult part wasn't remembering. It was knowing which memory still belonged to this world.

      after the change · main run

      No memory
      Shared memory

      Inheritance became inertia.

      Can JEV challenge its ancestors?

      one inherited Mitosis finding

      still · dodge · fire when ready

      inherited
      this world, so far
      status: ACTIVE

      after the change · main run · average points below best

      Memory only
      JEV + memory
      Simple rule + memory

      The rule beat JEV.

      Good.

      The experiment was supposed to teach us something, not flatter the model.

      In this small world a hand-built change detector was better: over the whole run it was ahead in of paired trials. JEV's value was flexible typed judgment in about ms, not magic superiority over deterministic code.

      So we moved the swarm onto real machines.

      games in the validation run
      worker sandboxes created and destroyed on schedule
      unplanned worker deaths
      /sampled games replayed exactly on a laptop
      /official Pilot games re-derived in clean Tenki sandboxes

      before the change · paired trials on Tenki

      Without memory
      JEV + memory

      Memory still helped: better in of trials.

      after the change

      Without memory
      JEV + memory

      The original JEV recovery did not replicate.

      Why?

      The memory was there. Our retrieval path wasn't returning all of it.

      In of reads, at least one of the trial's own findings was missing. None of the findings that came back was out of date.

      Memory is not only what you store.

      It is what the agent can recover when it matters.

      We read memory with a search call that returns its top matches, over one feed shared by parallel trials. Under that load our adapter returned an incomplete slice, so JEV often couldn't see the outdated finding it needed to challenge: its first challenge of one came at round , against in the main run. We froze the experiment instead of repairing it after seeing the result.

      a surprisingly small principle

      Representation shapes the decision landscape.

      A model does not act on the world directly. It acts on the representation of the world it can see.

      a question we left with

      An ecology of wisdom?

      Memory

      What from the past remains available?

      Salience

      What deserves attention now?

      Agency

      Can the system revise what it inherited?

      Perfect memory is not perfect adaptation.
      Forgetting everything prevents learning.
      Obeying everything remembered prevents change.

      Perhaps the interesting problem isn't how much an agent can remember. It's how experience should keep influencing action as the world changes.

      the stack, as an ecology

      Tenki

      Experience needs somewhere to happen.

      Disposable workers. Parallel games. Independent replay. Reproducible execution.

      without it here → no ephemeral population, no independent remote replay

      Mitosis

      Experience needs somewhere to persist.

      Shared findings survive worker death. New generations inherit earlier experiments.

      without it → worker death erases what was learned

      JEV

      Inheritance needs judgment.

      Fast typed decisions: what to do now, and which inherited findings deserve another look.

      without it → no learned, typed governance layer

      Remove any layer and you get a different system.

      Tenki is this implementation's execution substrate. The algorithm itself could run elsewhere; this swarm ran there.

      receipts

      Everything above, with its source.

      We started by asking whether JEV could play Space Invaders faster than an LLM.

      We ended up asking what happens when agents inherit experience in a world that refuses to stay still.

      Self-improvement isn't remembering everything.

      It's getting better at knowing what still applies.

      Play.Remember.Question.Adapt.