- Can a fast model play Space Invaders? Yes.
- Can experience make it better? Yes.
- Can what it learned become wrong? Yes.
- Can it know when to let go?
That's where things got interesting.
Memory helps.Until the world changes.
stable world · average points below the best setup · lower is better
The memory didn't disappear.
The world it described did.
First, we had to survive one frame.
Space Invaders made latency visible.
one second of play, slowed down · median per decision, holdout games
× faster
A language model could reason about the frame.
JEV could decide while the frame still mattered.
average score · untouched seeds, run once
show receipts
Then the benchmark started arguing back.
one real frame · what the model was shown
Same model. Same game.
Different representation.
Different policy.
So we tried something embarrassing:
a dumb rule.
sweep bot (no model) vs JEV on the compressed view · 50 fresh seeds · points without the bonus ship
High score was not enough evidence of understanding.
every strong result got a cheaper control
What survived: tell JEV which moves are safe, and to sweep unless it has a reason not to. On the holdout that player scored against the sweep bot's , and won seeds.
Play
What should I do now?
ms per moveEval
Did that decision survive reality?
on unseen life-or-death momentsLab
What deserves another experiment?
which setup gets the next games, and which old finding to doubtOne decision model.
Three timescales.
Let experience outlive the agent.
Four workers search ways to play: stand still, track, or sweep · dodge or not · fire when ready, or aimed. Each round, each worker plays one real game.
rounds later, we killed the workers.
The new workers had never played those games.
But the swarm had.
Tenki gave the swarm bodies.
Mitosis gave it inheritance.
This picture is the distributed run: real Tenki sandboxes. workers created, replaced exactly on schedule, unplanned deaths. Every newborn worker was checked at birth: it knew nothing. Findings shown are real inherited findings from trial 502.
stable world · main run · paired trials · pre-registered
average points below the best setup · lower is better
Shared memory cut search error by more than half.
- a new worker picks a strategy
- checks the shared inheritance
- skips the experiment a dead worker already ran
The swarm stopped paying twice for knowledge it had already earned.
Then we changed the world.
old world · true mean score of each setup
Some inheritance survived. Some became wrong.
The difficult part wasn't remembering. It was knowing which memory still belonged to this world.
after the change · main run
Inheritance became inertia.
Can JEV challenge its ancestors?
one inherited Mitosis finding
still · dodge · fire when ready
after the change · main run · average points below best
The rule beat JEV.
Good.
The experiment was supposed to teach us something, not flatter the model.
In this small world a hand-built change detector was better: over the whole run it was ahead in of paired trials. JEV's value was flexible typed judgment in about ms, not magic superiority over deterministic code.
So we moved the swarm onto real machines.
before the change · paired trials on Tenki
Memory still helped: better in of trials.
after the change
The original JEV recovery did not replicate.
Why?
The memory was there. Our retrieval path wasn't returning all of it.
In of reads, at least one of the trial's own findings was missing. None of the findings that came back was out of date.
Memory is not only what you store.
It is what the agent can recover when it matters.
We read memory with a search call that returns its top matches, over one feed shared by parallel trials. Under that load our adapter returned an incomplete slice, so JEV often couldn't see the outdated finding it needed to challenge: its first challenge of one came at round , against in the main run. We froze the experiment instead of repairing it after seeing the result.
a surprisingly small principle
- raw state→different policyfire jitter, points vs
- compressed state→different policysteady motion; with safe moves,
- full memory→different governancefirst outdated finding challenged at round
- partial retrieval→different governance…at round
- option order→different borderline judgment flips on borderline states vs on a plain repeat
Representation shapes the decision landscape.
A model does not act on the world directly. It acts on the representation of the world it can see.
a question we left with
An ecology of wisdom?
Memory
What from the past remains available?
Salience
What deserves attention now?
Agency
Can the system revise what it inherited?
Perfect memory is not perfect adaptation.
Forgetting everything prevents learning.
Obeying everything remembered prevents change.
Perhaps the interesting problem isn't how much an agent can remember. It's how experience should keep influencing action as the world changes.
the stack, as an ecology
Tenki
Experience needs somewhere to happen.
Disposable workers. Parallel games. Independent replay. Reproducible execution.
without it here → no ephemeral population, no independent remote replay
Mitosis
Experience needs somewhere to persist.
Shared findings survive worker death. New generations inherit earlier experiments.
without it → worker death erases what was learned
JEV
Inheritance needs judgment.
Fast typed decisions: what to do now, and which inherited findings deserve another look.
without it → no learned, typed governance layer
Remove any layer and you get a different system.
Tenki is this implementation's execution substrate. The algorithm itself could run elsewhere; this swarm ran there.
receipts
Everything above, with its source.
We started by asking whether JEV could play Space Invaders faster than an LLM.
We ended up asking what happens when agents inherit experience in a world that refuses to stay still.
Self-improvement isn't remembering everything.
It's getting better at knowing what still applies.
Play.Remember.Question.Adapt.