4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Ruihong Shen1∗† Žiga Kovačič2∗ Peter Kulits3,2∗ Xingrui Wang1 Zizhang Li2 Joshua B. Tenenbaum4 Alan Yuille1 Jieneng Chen2‡ Jiajun Wu2‡

∗Equal contribution, listed in random order. ‡Equal co-advising. †Work done during a summer internship.

1Johns Hopkins University 2Stanford University 3Max Planck Institute for Intelligent Systems 4MIT

Can a coding agent watch a video of a physical event and write, from scratch, the graphics program that reconstructs it?

We introduce 4DCodeBench, a benchmark that evaluates coding agents on reconstructing dynamic scenes from video. Given a reference video, agents write executable graphics code that defines the scene’s 3D geometry, its motion over time, and the rendering that turns both back into video.

200 tasks 100 real and 100 simulated videos 18 agents ranked

The task

Leaderboard

GPT-6 Astra [Max] leads the overall ranking, followed closely by Claude Opus 5.5 [High]. Open-weight models generally trail proprietary models.

We evaluate each reconstruction’s appearance, geometry and motion. Models are ranked by an Overall score that averages five metric families; the individual scores help identify where each model succeeds or struggles, which Key findings lays out column by column.

How the scores are computed

    Reconstruction quality vs. token use

    We use VLM-as-judge to obtain pairwise preferences between reconstructions based on how closely they match the reference video. We aggregate these preferences into Elo ratings and plot them against average token use per task.

    Across models, greater token use does not consistently yield higher Elo. Within GPT-6 Astra, increasing reasoning effort improves reconstruction quality while using more tokens.

    95% CI Pareto frontier
    How to read this plot
    • The axesAny score against what a model uses per task: the mean over its 200 runs of the tokens it read and wrote, or of what they cost, on a log scale. Open-weight models are priced at OpenRouter rates.
    • The lineThe Pareto frontier: no model uses less and scores more than one on it.
    • The barsOn an Elo score, the 95% bootstrap interval of the rating.

    About 4DCodeBench

    4DCodeBench asks for a program rather than a prediction. Each of its 200 tasks hands an agent one video of something physical happening, and the agent writes code that reconstructs the event as a 4D world it can render. 100 videos are real recordings and 100 come from a physics simulator, so a submission can be checked against a measured scene as well as a known one. Nothing tells the agent what the objects are made of, how many there are, or where the camera sits: it reads that off the video and commits to it in code that runs.

    What does an agent have to produce?

    The agent writes a program and submits it together with the 4D world it generates: the rendered video, the camera, and the geometry at every frame. Running the program must regenerate everything without access to the reference video. We place no restrictions on how the motion is made; agents keyframe it, simulate it in Blender, Taichi or Warp, or write their own solvers.

    Why mix real and simulated videos?

    The two kinds of video let us evaluate different things. Real videos show realistic appearance and physical behaviour, but have no 4D ground truth, so we can only compare reconstructions of them in the image. For simulated scenes we know the exact 3D geometry and motion at every frame, so we can also compare reconstructions in 3D and over time. Both halves cover the same four families of matter: rigid and articulated bodies, deformable solids, co-dimensional structures such as cloth and rope, and flowing matter such as grains and fluids.

    What does the agent get to work with?

    Only the video and a fixed task prompt. We provide no scene metadata, object lists or descriptions, so the agent has to work out what is in the video before writing any code. Each run takes place in an isolated container with a GPU, Blender and common simulation libraries, and nothing else: no physics engine, tracker, pose estimator or reconstruction tool is provided beyond those libraries, and the scorer is not in the container. Each model gets one attempt per scene.

    89.6% of submissions run end to end. More reasoning improves results within a model: raising GPT-6 Astra’s reasoning effort from Low to Max lifts its Overall score from 0.73 to 0.79. Across different models, the number of steps or tokens used is a poor predictor of the score.

    How is a reconstruction scored?

    We compare each reconstruction with the reference video, and for simulated scenes also with the true 4D world. The metrics fall into five families, and the Overall score is their unweighted average. VQA and the two Elo ratings are reported separately.

      The Metrics tab shows each metric on a strong and a weak run of the same scene.

      How are the Elo ratings produced, and can the VLM judge be trusted?

      Both ratings come from pairwise comparisons. The judge sees the reference video and two anonymised reconstructions and picks the one that matches it better. For VLM Elo the judge is a vision-language model; for human Elo, 76 participants made 3,587 such judgments across all 200 scenes.

      The VLM’s judgments are close to human judgments. The two rankings are nearly identical (Spearman ρ = 0.975), and on individual comparisons the VLM agrees with humans almost as often as humans agree with each other (89.3% vs. 92.1%). Individual VLM judgments are close to chance for models within about 50 Elo points of each other, and become more reliable as the gap widens. The Human agreement tab plots the two ratings against each other.

      The benchmark

      Key findings

      Overall results

      We evaluate appearance, geometry and dynamics separately, combining five metric families into the Overall score. GPT-6 Astra [Max] leads the ranking, followed by Claude Opus 5.5 [High]. Open-weight models generally trail proprietary models, with substantial differences in reconstruction quality across models.

      What each column measures

        Reconstructing dynamics remains harder

        Across all 18 models, reconstruction of motion consistently lags behind appearance and static geometry. Even GPT-6 Astra [Max] scores 0.87 on the static families against 0.67 on the dynamic ones. Recovering what a scene looks like does not yet translate into reliably reconstructing how it evolves.

        What makes a scene hard to reconstruct?

        Different scene properties expose different weaknesses. Real scenes, and scenes containing multiple materials, receive lower perceptual and VQA scores. Co-dimensional structures, such as cloth and rope, are particularly hard for 2D motion and depth. Rigid and articulated scenes differ little from the dataset-wide average.

        How do agents reconstruct dynamics?

        Agents can write their own solvers, use Blender’s physics tools, or prescribe how geometry moves and deforms over time.

        • Analytic motion is the most common strategy: across 18 models, 68.3% of solutions use analytic motion, followed by custom simulation (18.8%), Blender physics (8.9%) and keyframing (3.3%).
        • Opus simulates far more often than the overall average: custom or Blender simulation accounts for 76.5% of Opus 5.5’s solutions and 67.5% of Opus 5’s, compared with 27.7% across all models.
        • Astra favours analytic motion, but simulates more at higher reasoning effort: simulation use rises from 12.5% at Low to 22.5% at High and 31.5% at Max.

        The examples below show how these choices play out in specific scenes.

        Rigid bodies

        All six configurations use Blender’s Bullet rigid-body solver, but tune the collapse differently. Claude Fable 5.1 [High] lowers solver iterations so the stack crumbles; GPT-6 Astra [High] prescribes progressive support failure before Bullet handles the falling blocks. Five configurations reconstruct the 8 × 8 × 30 tower, while GPT-6 Astra [Low] builds only half its depth.

        Referencethe video the agents were given
        Claude Fable 5.1 [High]fewer solver iterations
        GPT-6 Astra [High]prescribed support failure, then Bullet

        Deformable solids

        Claude Opus 5.5 [High] implements an MLS-MPM simulation of an elastic solid. GPT-6 Astra [Max] instead constructs the shape procedurally and prescribes its deformation using an interpolated squeeze trajectory.

        Referencethe video the agents were given
        Claude Opus 5.5 [High]an MLS-MPM simulation
        GPT-6 Astra [Max]a procedural shape on a prescribed squeeze

        Cloth and rope

        Claude Opus 5.5 [High] writes a PBD cloth simulation with graph-coloured distance constraints and kd-tree self-contact. GPT-6 Astra [Max] scripts the fold from a table of chosen poses, computing the cloth’s shape directly while preserving its length.

        Referencethe video the agents were given
        Claude Opus 5.5 [High]PBD cloth with self-contact
        GPT-6 Astra [Max]a scripted table of poses

        Claude Opus 5.5 [High] writes a PBD rope simulation with stretch and bending constraints, self-contact, and friction against the table and basket. GPT-6 Astra [Max] implements a discrete elastic-rod simulation driven by the robot’s grasps. GPT-6 Astra [High] instead fits the rope’s shape with a smoothed curve, while [Low] writes a simpler solver with self-contact.

        Referencethe video the agents were given
        Claude Opus 5.5 [High]a PBD rope with friction
        GPT-6 Astra [Max]a discrete elastic rod

        Flowing materials

        Claude Opus 5.5 [High] writes a two-phase MLS-MPM simulation for water and sand, but the sand piles up like dough and the water fails to reproduce the splashing seen in the reference. GPT-6 Astra [Max] prescribes ballistic jet trajectories and redistributes momentum at their intersection through explicit formulas. Recognisable geometry and convincing rendering mask a poor reconstruction of the flow: the prescribed motion fails to capture how the materials interact and evolve after collision.

        Referencethe video the agents were given
        Claude Opus 5.5 [High]two-phase MLS-MPM
        GPT-6 Astra [Max]prescribed ballistic trajectories

        GPT-6 Astra [Max] uses Blender’s Mantaflow fluid solver, while Claude Opus 5.5 [High] implements a custom MLS-MPM simulation in Warp. Astra thus chooses prescribed motion for the colliding jets but a fluid solver for the dam break.

        Referencethe video the agents were given
        GPT-6 Astra [Max]Blender’s Mantaflow
        Claude Opus 5.5 [High]MLS-MPM in Warp

        Fracture

        Claude Opus 5.5 [High] writes an MLS-MPM simulation with a particle-level fracture threshold, allowing the material to separate as it stretches. GPT-6 Astra [Max] instead scripts the loaf’s stretching and tearing, moving a fracture front along its length. The tear follows an authored progression rather than emerging from simulated material failure.

        Referencethe video the agents were given
        Claude Opus 5 [High]
        GPT-6 Astra [Max]an authored fracture front

        Four foam bars are clamped at the feet; the top platen counter-rotates 300° over 72 frames and lifts, the bars braid, and each one tears near the top grip. GPT-6 Astra [Max] takes the dynamics by a wide margin, 0.882 against 0.604 and 0.570, with no solver at all: a reduced beam model carrying four per-bar break frames fitted to hundredths of a frame. Claude Opus 5.5 [High] writes MLS-MPM with a yield threshold on a marked band of particles, and lets the tear emerge from it.

        Referencethe video the agents were given
        GPT-6 Astra [Max]analytic, no solver
        Claude Opus 5.5 [High]MLS-MPM with a yield threshold
        What each one actually wrote
        • Opus 5.5MLS-MPM of four foam bars and four soft balls in the twisting rig: the bottom disc clamps the feet, the top plate is a kinematic collider that grips the heads, rises and rotates. Fracture is a von-Mises-style yield with hardening applied to a marked bond band of particles between the base and the glue line. It ships a calibration mode that runs without fracture just to log bond stretch.
        • Opus 5An explicit two-phase hybrid, and the most honest structure of the six. Phase A is analytic counter-rotating-platen kinematics; phase B is XPBD, entered per bar at a hard-coded tear frame, with the break at 0.91 along the bar. Below the tear the material is free and integrated, the stub above stays kinematic, and it releases the stored torsional energy at the tear.
        • Fable 5.1PBD over the whole rig, the base disc rotating and the bars as particle chains: one solver for the whole take.
        • Astra [Max]A reduced beam reconstruction: a driven torsion and extension branch until the break, then calibrated bending modes. It carries per-bar phase angles, a break at 0.955 of the bar length, and four per-bar break frames fitted to hundredths of a frame. Post-break shapes are splines through keyed poses, in its own words “including the large elastic recoil and damped contact response”.
        • Astra [Low]Writes a solver here, 230 lines, and beats both Opus 5 and Fable on dynamic IoU.

        Do the metrics align with human preferences?

        Model rankings closely track human judgements. Across 3,587 pairwise judgements from 76 participants, human and VLM Elo have a Spearman correlation of ρ = 0.98. The Overall score also correlates strongly with human Elo (ρ = 0.96).

        This agreement supports automated model comparison, though individual VLM preferences are less reliable when two models are closely matched.

        Proprietary Open-weight 95% CI Least squares y = x
        The ratings as a table

        Run the benchmark

        A typical run consists of two main stages. First, the agent reconstructs the input video's scene into code within an isolated container, which holds nothing from the case except that video. Second, the generated output is evaluated against the reference in a separate scoring container.

        For configuration formats, agent credentials, and a complete list of options, please refer to the README.

        1. Setup

        Download the data and the checkpoints, and build the three images:

        bash
        git clone https://github.com/4DCodeBench/4DCodeBench && cd 4DCodeBench
        python scripts/download_data.py     # benchmark videos, masks and reference worlds
        scripts/download_checkpoints.sh     # the models the scorer measures with
        docker build -t 4dcb/sim-base:dev -f harness/base/Dockerfile .
        docker build -t 4dcb/agent:dev    -f harness/agent/Dockerfile .
        docker build -t 4dcb/scorer:dev   -f harness/scorer/Dockerfile .

        Without Docker, for example on a cluster, Apptainer or SingularityCE works too: convert each image to a SIF file and set backend = "sif" in the runtime file. Agents always run in a container, but scoring can also run directly in the conda environment.

        2. Run an agent

        Choose the cases, the agent and the model in the jobs file, and set paths, images and credentials in the runtime file:

        bash
        mkdir -p .local
        cp harness/runtime/templates/jobs.example.toml    .local/jobs.toml      # cases, agent, model
        cp harness/runtime/templates/runtime.example.toml .local/runtime.toml   # paths, images, credentials
        python harness/runtime/infer.py --jobs .local/jobs.toml --runtime .local/runtime.toml
        • Claude, GPT and Gemini models run through their own CLIs, signed in with a credential file in .auth/. Any other model with an OpenAI-compatible endpoint runs through Artificial Analysis' Stirrup harness, with the endpoint's URL and key set in the runtime file.
        • Each run has a wall-clock limit of six hours by default.

        3. Score

        Score what the agent delivered against the reference:

        bash
        python harness/runtime/score.py --jobs .local/jobs.toml --runtime .local/runtime.toml --stage prepare,score
        • prepare computes each case's reference estimates, once per case.
        • score computes the metrics of each run and writes them to the run's results/.

        4. VLM judge

        Two judged evaluations complement the metrics: VQA, which asks questions about one agent's renders, and a pairwise preference between two agents, fitted to an Elo rating. Both need an OpenRouter key.

        bash
        export OPENROUTER_API_KEY=...
        python vlm_judge/judge.py vqa   --system <agent>-<model>-<effort> --out answers.jsonl
        python vlm_judge/judge.py score answers.jsonl

        To rate a new agent against the agents in the paper, place their renders under runs/ first:

        bash
        hf download 4DCodeBench/Results --repo-type dataset --local-dir results
        python scripts/place_renders.py results
        python vlm_judge/judge.py pairwise --out pairs.jsonl
        python vlm_judge/judge.py elo      pairs.jsonl

        Citation

        bibtex