-
Where the Condition Enters
Five papers that between them define every current video generation model — DiT, Rectified Flow, SD3, ControlNet and Self-Forcing. Architecture, objective and main figure for each, organised around one question that turns out to decide a lot: when you have something to tell a video model, which of the five available doors do you push it through?
-
Thinking Between Frames
A survey of test-time optimization for robot policies — five families, from best-of-N with a learned verifier through world-model search to fast-weight updates during deployment. The organizing constraint is one nobody in LLM test-time scaling has to think about: a 30 Hz control loop gives you 33 milliseconds, and a single policy forward pass already costs 73.
-
The Proxy Is the Contract
Paper notes on Code World Model. A coding agent maintains world state as executable Python, a deterministic compiler rasterises it into a coarse proxy video, and a video model renders that into pixels. The architecture is the one I would have drawn — and the paper ships zero benchmarks, which makes the interesting question not whether it works but which of its two load-bearing choices anyone has actually tested.
-
The Exchange Rate
Perceptron open-sourced Isaac 0.5, a 36B sparse embodied foundation model, and with it the first published price list for trading cheap video against expensive teleoperation. The headline is 210×. The technical report contains three caveats that move that number to somewhere between 5.7× and 300×, and a $1M build plan that says something the headline does not.
-
Retrieval Is Not Memory
How memory is actually implemented in video world models — five mechanisms, all of them attention over past frames — how long it lasts in practice, and the structural reason none of it can know what happened while the camera was pointed elsewhere.
-
Four Places to Put the Code
A survey of the LLM-writes-code-for-robots lineage, from Code as Policies in 2022 to Code World Model this week. The organizing observation is that the code has been steadily retreating from the actuator toward the world — first the policy, then the objective, then the domain, now the state — and every step was forced by the same failure.
-
Five Papers Under Every Video Model
The architecture and objective every current video model shares, read through the five papers that fixed it — and the one question they answer that the interactive-world debate is still arguing about: where conditioning enters, and what that costs.
-
The Frame Is Not the World
Loopit open-sourced a real-time interactive video world model and then argued its own approach cannot get to the endgame — code has to hold world state, pixels only render it. I read the release, checked the leaderboard claim, and found the benchmark had already published the evidence for their thesis without connecting it.
-
Show, Don't Tell
Paper notes on Masked Visual Actions for Unified World Modeling. The contribution is not a world model — it is an interface: express a robot action as a partially revealed pixel trajectory, and one LoRA on an off-the-shelf video model becomes a forward model, an inverse model, a policy evaluator and a planner.
-
Two Substrates of Self-Evolving Agents
Thirty-odd papers, one entry each, with the architecture figure from the original. One school writes improvement into weights; the other writes it into text. What decides whether either works is neither — it is the gap between generating an answer and checking one.