4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
∗Equal contribution, listed in random order. ‡Equal co-advising. †Work done during a summer internship.
1
2
3
4
Paper Citation Leaderboard Data Code
Can a coding agent watch a video of a physical event and write, from scratch, the graphics program that reconstructs it?
We introduce 4DCodeBench, a benchmark that evaluates coding agents on reconstructing dynamic scenes from video. Given a reference video, agents write executable graphics code that defines the scene’s 3D geometry, its motion over time, and the rendering that turns both back into video.
The task
Leaderboard
GPT-6 Astra [Max] leads the overall ranking,
followed closely by
Claude Opus 5.5 [High].
Open-weight models generally trail proprietary models.
We evaluate each reconstruction’s appearance, geometry and motion. Models are ranked by an Overall score that averages five metric families; the individual scores help identify where each model succeeds or struggles, which Key findings lays out column by column.
How the scores are computed
Reconstruction quality vs. token use
We use VLM-as-judge to obtain pairwise preferences between reconstructions based on how closely they match the reference video. We aggregate these preferences into Elo ratings and plot them against average token use per task.
Across models, greater token use does not consistently yield
higher Elo. Within GPT-6 Astra, increasing reasoning effort improves
reconstruction quality while using more tokens.
How to read this plot
- The axesA score against what a model spends per task, averaged over its 200 runs: tokens read and written, or their cost, on a log scale. Open-weight models are priced at OpenRouter rates.
- The lineThe Pareto frontier: no other model both spends less and scores higher than a model on it.
- The barsFor Elo scores, the 95% bootstrap interval.
- Real / SyntheticWith one split selected, Elo, VQA and usage are recomputed on its 100 tasks alone.
About 4DCodeBench
4DCodeBench asks for a program rather than a prediction. Each of its 200 tasks hands an agent one video of something physical happening, and the agent writes code that reconstructs the event as a 4D world it can render. 100 videos are real recordings and 100 come from a physics simulator, so a submission can be checked against a measured scene as well as a known one. Nothing tells the agent what the objects are made of, how many there are, or where the camera sits: it reads that off the video and commits to it in code that runs.
What does an agent have to produce?
The agent writes a program and submits it together with the 4D world it generates: the rendered video, the camera, and the geometry at every frame. Running the program must regenerate everything without access to the reference video. We place no restrictions on how the motion is made; agents keyframe it, simulate it in Blender, Taichi or Warp, or write their own solvers.
Why mix real and simulated videos?
The two kinds of video let us evaluate different things. Real videos show realistic appearance and physical behaviour, but have no 4D ground truth, so we can only compare reconstructions of them in the image. For simulated scenes we know the exact 3D geometry and motion at every frame, so we can also compare reconstructions in 3D and over time. Both halves cover the same four families of matter: rigid and articulated bodies, deformable solids, co-dimensional structures such as cloth and rope, and flowing matter such as grains and fluids.
What does the agent get to work with?
Only the video and a fixed task prompt. We provide no scene metadata, object lists or descriptions, so the agent has to work out what is in the video before writing any code. Each run takes place in an isolated container with a GPU, Blender and common simulation libraries, and nothing else: no physics engine, tracker, pose estimator or reconstruction tool is provided beyond those libraries, and the scorer is not in the container. Each model gets one attempt per scene.
90.1% of submissions run end to end. More reasoning improves results within a model: raising GPT-6 Astra’s reasoning effort from Low to Max lifts its Overall score from 0.73 to 0.79. Across different models, the number of steps or tokens used is a poor predictor of the score.
How is a reconstruction scored?
We compare each reconstruction with the reference video, and for simulated scenes also with the true 4D world. The metrics fall into five families, and the Overall score is their unweighted average. VQA and the two Elo ratings are reported separately.
The Metrics tab shows each metric on a strong and a weak run of the same scene.
How are the Elo ratings produced, and can the VLM judge be trusted?
Both ratings come from pairwise comparisons. The judge sees the reference video and two anonymised reconstructions and picks the one that matches it better. For VLM Elo the judge is a vision-language model; for human Elo, 76 participants made 3,587 such judgments across all 200 scenes.
The VLM’s judgments are close to human judgments. The two rankings are nearly identical (Spearman ρ = 0.98), and on individual comparisons the VLM agrees with humans almost as often as humans agree with each other (89.3% vs. 92.1%). Individual VLM judgments are close to chance for models within about 50 Elo points of each other, and become more reliable as the gap widens. The Human agreement tab plots the two ratings against each other.
How do I run my own model on the benchmark?
Running a model takes four steps. The agent works in an isolated container that holds only the input video; scoring runs in a separate one.
- 1. SetupDownload the data and checkpoints and build the images, with Docker or, on a cluster, Apptainer or SingularityCE.
- 2. Run an agentChoose the cases, agent and model in the jobs file. Claude, GPT and Gemini run through their own CLIs; any model with an OpenAI-compatible endpoint runs through Stirrup.
- 3. ScoreCompute each case’s reference estimates once, then the metrics of every run.
- 4. VLM judgeVQA and the pairwise Elo, with an OpenRouter key. Download our agents’ renders to rate a new model against them.
Commands, configuration formats and credentials are in the README.
The benchmark
Input video
Agent reconstructions
Drag across any render to compareWhat materials are present in the scenes realsimulated
Click any task to see every model’s result
Key findings
Overall results
We evaluate appearance, geometry and dynamics separately, combining five metric
families into the Overall score. GPT-6 Astra [Max] leads the ranking,
followed by
Claude Opus 5.5 [High]. Open-weight models generally trail
proprietary models, with substantial differences in reconstruction quality across
models.
What each column measures
Reconstructing dynamics remains harder
Across all 18 models, reconstruction of motion consistently lags behind
appearance and static geometry. Even GPT-6 Astra [Max] scores
0.91 on the static families against 0.67 on the dynamic ones.
Recovering what a scene looks like does not yet translate into reliably
reconstructing how it evolves.
What makes a scene hard to reconstruct?
Different scene properties expose different weaknesses. Real scenes, and scenes containing multiple materials, receive lower perceptual and VQA scores. Co-dimensional structures, such as cloth and rope, are particularly hard for 2D motion and depth. Rigid and articulated scenes differ little from the dataset-wide average.
How do agents reconstruct dynamics?
Agents can write their own solvers, use Blender’s physics tools, or prescribe how geometry moves and deforms over time.
- Analytic motion is the most common strategy: across 18 models, 67% of solutions use analytic motion, followed by custom simulation (19%), Blender physics (10%) and keyframing (3%).
- Opus simulates far more often than the overall average: custom or Blender simulation accounts for 76% of Opus 5.5’s solutions and 67% of Opus 5’s, compared with 29% across all models.
- Astra favours analytic motion, but simulates more at higher reasoning effort: simulation use rises from 13% at Low to 22% at High and 32% at Max.
The examples below show how these choices play out in specific scenes.
Rigid bodies
All six configurations use Blender’s Bullet rigid-body solver, but tune the collapse differently. Claude Fable 5.1 [High] lowers solver iterations so the stack crumbles;
GPT-6 Astra [High] prescribes progressive support failure before Bullet handles the falling blocks. Five configurations reconstruct the 8 × 8 × 30 tower, while Astra [Low] builds only half its depth.
Deformable solids
Claude Opus 5.5 [High] implements an MLS-MPM simulation of an elastic solid.
GPT-6 Astra [Max] instead constructs the shape procedurally and prescribes its deformation using an interpolated squeeze trajectory.
Cloth and rope
Claude Opus 5.5 [High] writes a PBD cloth simulation with graph-coloured distance constraints and kd-tree self-contact.
GPT-6 Astra [Max] scripts the fold from a table of chosen poses, computing the cloth’s shape directly while preserving its length.
Claude Opus 5.5 [High] writes a PBD rope simulation with stretch and bending constraints, self-contact, and friction against the table and basket.
GPT-6 Astra [Max] implements a discrete elastic-rod simulation driven by the robot’s grasps. Astra [High] instead fits the rope’s shape with a smoothed curve, while Astra [Low] writes a simpler solver with self-contact.
Flowing materials
Claude Opus 5.5 [High] writes a two-phase MLS-MPM simulation for water and sand, but the sand piles up like dough and the water fails to reproduce the splashing seen in the reference.
GPT-6 Astra [Max] prescribes ballistic jet trajectories and redistributes momentum at their intersection through explicit formulas. Recognisable geometry and convincing rendering mask a poor reconstruction of the flow: the prescribed motion fails to capture how the materials interact and evolve after collision.
For the dam break, GPT-6 Astra [Max] uses Blender’s Mantaflow fluid solver, while
Claude Opus 5.5 [High] implements a custom MLS-MPM simulation in taichi. Models switch strategies as the dynamics grow more complex.
Fracture
Claude Opus 5.5 [High] writes an MLS-MPM simulation with a particle-level fracture threshold, allowing the material to separate as it stretches.
GPT-6 Astra [Max] instead scripts the loaf’s stretching and tearing, moving a fracture front along its length. The tear follows an authored progression rather than emerging from simulated material failure.
Four foam bars are clamped at the feet; the top platen counter-rotates
300° over 72 frames and lifts, the bars braid, and each one tears near
the top grip. GPT-6 Astra [Max] leads on dynamics by 31% over the
next-best model, with no solver at all: a reduced beam model
carrying four per-bar break frames fitted to hundredths of a frame.
Claude Opus 5.5 [High]
writes MLS-MPM with a yield threshold on a marked band of particles, and lets
the tear emerge from it.
Do the metrics align with human preferences?
Model rankings closely track human judgements. Across 3,587 pairwise judgements from 76 participants, human and VLM Elo have a Spearman correlation of ρ = 0.98. The Overall score also correlates strongly with human Elo (ρ = 0.96).
This agreement supports automated model comparison, though individual VLM preferences are less reliable when two models are closely matched.
The ratings as a table
Citation