4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Ruihong Shen1∗† Žiga Kovačič2∗ Peter Kulits3,2∗ Xingrui Wang1 Zizhang Li2 Joshua B. Tenenbaum4 Alan Yuille1 Jieneng Chen2‡ Jiajun Wu2‡

∗Equal contribution, listed in random order. ‡Equal co-advising. †Work done during a summer internship.

1Johns Hopkins University 2Stanford University 3Max Planck Institute for Intelligent Systems 4MIT

Can a coding agent watch a video of a physical event and write, from scratch, the graphics program that reconstructs it?

We introduce 4DCodeBench, a benchmark that evaluates coding agents on reconstructing dynamic scenes from video. Given a reference video, agents write executable graphics code that defines the scene’s 3D geometry, its motion over time, and the rendering that turns both back into video.

200 tasks 100 real and 100 simulated videos 18 agents ranked

The task

Leaderboard

GPT-6 Astra [Max] leads the overall ranking, followed closely by Claude Opus 5.5 [High].
Open-weight models generally trail proprietary models.

We evaluate each reconstruction’s appearance, geometry and motion. Models are ranked by an Overall score that averages five metric families; the individual scores help identify where each model succeeds or struggles, which Key findings lays out column by column.

How the scores are computed

    Reconstruction quality vs. token use

    We use VLM-as-judge to obtain pairwise preferences between reconstructions based on how closely they match the reference video. We aggregate these preferences into Elo ratings and plot them against average token use per task.

    Across models, greater token use does not consistently yield higher Elo. Within GPT-6 Astra, increasing reasoning effort improves reconstruction quality while using more tokens.

    95% CI Pareto frontier
    How to read this plot
    • The axesA score against what a model spends per task, averaged over its 200 runs: tokens read and written, or their cost, on a log scale. Open-weight models are priced at OpenRouter rates.
    • The lineThe Pareto frontier: no other model both spends less and scores higher than a model on it.
    • The barsFor Elo scores, the 95% bootstrap interval.
    • Real / SyntheticWith one split selected, Elo, VQA and usage are recomputed on its 100 tasks alone.

    About 4DCodeBench

    4DCodeBench asks for a program rather than a prediction. Each of its 200 tasks hands an agent one video of something physical happening, and the agent writes code that reconstructs the event as a 4D world it can render. 100 videos are real recordings and 100 come from a physics simulator, so a submission can be checked against a measured scene as well as a known one. Nothing tells the agent what the objects are made of, how many there are, or where the camera sits: it reads that off the video and commits to it in code that runs.

    What does an agent have to produce?

    The agent writes a program and submits it together with the 4D world it generates: the rendered video, the camera, and the geometry at every frame. Running the program must regenerate everything without access to the reference video. We place no restrictions on how the motion is made; agents keyframe it, simulate it in Blender, Taichi or Warp, or write their own solvers.

    Why mix real and simulated videos?

    The two kinds of video let us evaluate different things. Real videos show realistic appearance and physical behaviour, but have no 4D ground truth, so we can only compare reconstructions of them in the image. For simulated scenes we know the exact 3D geometry and motion at every frame, so we can also compare reconstructions in 3D and over time. Both halves cover the same four families of matter: rigid and articulated bodies, deformable solids, co-dimensional structures such as cloth and rope, and flowing matter such as grains and fluids.

    What does the agent get to work with?

    Only the video and a fixed task prompt. We provide no scene metadata, object lists or descriptions, so the agent has to work out what is in the video before writing any code. Each run takes place in an isolated container with a GPU, Blender and common simulation libraries, and nothing else: no physics engine, tracker, pose estimator or reconstruction tool is provided beyond those libraries, and the scorer is not in the container. Each model gets one attempt per scene.

    90.1% of submissions run end to end. More reasoning improves results within a model: raising GPT-6 Astra’s reasoning effort from Low to Max lifts its Overall score from 0.73 to 0.79. Across different models, the number of steps or tokens used is a poor predictor of the score.

    How is a reconstruction scored?

    We compare each reconstruction with the reference video, and for simulated scenes also with the true 4D world. The metrics fall into five families, and the Overall score is their unweighted average. VQA and the two Elo ratings are reported separately.

      The Metrics tab shows each metric on a strong and a weak run of the same scene.

      How are the Elo ratings produced, and can the VLM judge be trusted?

      Both ratings come from pairwise comparisons. The judge sees the reference video and two anonymised reconstructions and picks the one that matches it better. For VLM Elo the judge is a vision-language model; for human Elo, 76 participants made 3,587 such judgments across all 200 scenes.

      The VLM’s judgments are close to human judgments. The two rankings are nearly identical (Spearman ρ = 0.98), and on individual comparisons the VLM agrees with humans almost as often as humans agree with each other (89.3% vs. 92.1%). Individual VLM judgments are close to chance for models within about 50 Elo points of each other, and become more reliable as the gap widens. The Human agreement tab plots the two ratings against each other.

      How do I run my own model on the benchmark?

      Running a model takes four steps. The agent works in an isolated container that holds only the input video; scoring runs in a separate one.

      • 1. SetupDownload the data and checkpoints and build the images, with Docker or, on a cluster, Apptainer or SingularityCE.
      • 2. Run an agentChoose the cases, agent and model in the jobs file. Claude, GPT and Gemini run through their own CLIs; any model with an OpenAI-compatible endpoint runs through Stirrup.
      • 3. ScoreCompute each case’s reference estimates once, then the metrics of every run.
      • 4. VLM judgeVQA and the pairwise Elo, with an OpenRouter key. Download our agents’ renders to rate a new model against them.

      Commands, configuration formats and credentials are in the README.

      The benchmark

      Key findings

      Overall results

      We evaluate appearance, geometry and dynamics separately, combining five metric families into the Overall score. GPT-6 Astra [Max] leads the ranking, followed by Claude Opus 5.5 [High]. Open-weight models generally trail proprietary models, with substantial differences in reconstruction quality across models.

      What each column measures

        Reconstructing dynamics remains harder

        Across all 18 models, reconstruction of motion consistently lags behind appearance and static geometry. Even GPT-6 Astra [Max] scores 0.91 on the static families against 0.67 on the dynamic ones. Recovering what a scene looks like does not yet translate into reliably reconstructing how it evolves.

        What makes a scene hard to reconstruct?

        Different scene properties expose different weaknesses. Real scenes, and scenes containing multiple materials, receive lower perceptual and VQA scores. Co-dimensional structures, such as cloth and rope, are particularly hard for 2D motion and depth. Rigid and articulated scenes differ little from the dataset-wide average.

        How do agents reconstruct dynamics?

        Agents can write their own solvers, use Blender’s physics tools, or prescribe how geometry moves and deforms over time.

        • Analytic motion is the most common strategy: across 18 models, 67% of solutions use analytic motion, followed by custom simulation (19%), Blender physics (10%) and keyframing (3%).
        • Opus simulates far more often than the overall average: custom or Blender simulation accounts for 76% of Opus 5.5’s solutions and 67% of Opus 5’s, compared with 29% across all models.
        • Astra favours analytic motion, but simulates more at higher reasoning effort: simulation use rises from 13% at Low to 22% at High and 32% at Max.

        The examples below show how these choices play out in specific scenes.

        Rigid bodies

        All six configurations use Blender’s Bullet rigid-body solver, but tune the collapse differently. Claude Fable 5.1 [High] lowers solver iterations so the stack crumbles; GPT-6 Astra [High] prescribes progressive support failure before Bullet handles the falling blocks. Five configurations reconstruct the 8 × 8 × 30 tower, while Astra [Low] builds only half its depth.

        Referencethe video the agents were given
        Claude Fable 5.1 [High]fewer solver iterations
        GPT-6 Astra [High]prescribed support failure, then Bullet
        GPT-6 Astra [Low]half the tower’s depth

        Deformable solids

        Claude Opus 5.5 [High] implements an MLS-MPM simulation of an elastic solid. GPT-6 Astra [Max] instead constructs the shape procedurally and prescribes its deformation using an interpolated squeeze trajectory.

        Referencethe video the agents were given
        Claude Opus 5.5 [High]an MLS-MPM simulation
        GPT-6 Astra [Max]a procedural shape on a prescribed squeeze

        Cloth and rope

        Claude Opus 5.5 [High] writes a PBD cloth simulation with graph-coloured distance constraints and kd-tree self-contact. GPT-6 Astra [Max] scripts the fold from a table of chosen poses, computing the cloth’s shape directly while preserving its length.

        Referencethe video the agents were given
        Claude Opus 5.5 [High]PBD cloth with self-contact
        GPT-6 Astra [Max]a scripted table of poses

        Claude Opus 5.5 [High] writes a PBD rope simulation with stretch and bending constraints, self-contact, and friction against the table and basket. GPT-6 Astra [Max] implements a discrete elastic-rod simulation driven by the robot’s grasps. Astra [High] instead fits the rope’s shape with a smoothed curve, while Astra [Low] writes a simpler solver with self-contact.

        Referencethe video the agents were given
        Claude Opus 5.5 [High]a PBD rope with friction
        GPT-6 Astra [Max]a discrete elastic rod
        GPT-6 Astra [High]a smoothed curve
        GPT-6 Astra [Low]a simpler solver with self-contact

        Flowing materials

        Claude Opus 5.5 [High] writes a two-phase MLS-MPM simulation for water and sand, but the sand piles up like dough and the water fails to reproduce the splashing seen in the reference. GPT-6 Astra [Max] prescribes ballistic jet trajectories and redistributes momentum at their intersection through explicit formulas. Recognisable geometry and convincing rendering mask a poor reconstruction of the flow: the prescribed motion fails to capture how the materials interact and evolve after collision.

        Referencethe video the agents were given
        Claude Opus 5.5 [High]two-phase MLS-MPM
        GPT-6 Astra [Max]prescribed ballistic trajectories

        For the dam break, GPT-6 Astra [Max] uses Blender’s Mantaflow fluid solver, while Claude Opus 5.5 [High] implements a custom MLS-MPM simulation in taichi. Models switch strategies as the dynamics grow more complex.

        Referencethe video the agents were given
        GPT-6 Astra [Max]Blender’s Mantaflow
        Claude Opus 5.5 [High]MLS-MPM in taichi

        Fracture

        Claude Opus 5.5 [High] writes an MLS-MPM simulation with a particle-level fracture threshold, allowing the material to separate as it stretches. GPT-6 Astra [Max] instead scripts the loaf’s stretching and tearing, moving a fracture front along its length. The tear follows an authored progression rather than emerging from simulated material failure.

        Referencethe video the agents were given
        Claude Opus 5.5 [High]
        GPT-6 Astra [Max]an authored fracture front

        Four foam bars are clamped at the feet; the top platen counter-rotates 300° over 72 frames and lifts, the bars braid, and each one tears near the top grip. GPT-6 Astra [Max] leads on dynamics by 31% over the next-best model, with no solver at all: a reduced beam model carrying four per-bar break frames fitted to hundredths of a frame. Claude Opus 5.5 [High] writes MLS-MPM with a yield threshold on a marked band of particles, and lets the tear emerge from it.

        Referencethe video the agents were given
        GPT-6 Astra [Max]analytic, no solver
        Claude Opus 5.5 [High]MLS-MPM with a yield threshold

        Do the metrics align with human preferences?

        Model rankings closely track human judgements. Across 3,587 pairwise judgements from 76 participants, human and VLM Elo have a Spearman correlation of ρ = 0.98. The Overall score also correlates strongly with human Elo (ρ = 0.96).

        This agreement supports automated model comparison, though individual VLM preferences are less reliable when two models are closely matched.

        Proprietary Open-weight 95% CI Least squares y = x
        The ratings as a table

        Citation

        bibtex