4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

Ruihong Shen1∗† Žiga Kovačič2∗ Peter Kulits3,2∗ Xingrui Wang1 Zizhang Li2 Joshua B. Tenenbaum4 Alan Yuille1 Jieneng Chen2‡ Jiajun Wu2‡

∗Equal contribution, listed in random order. ‡Equal co-advising. †Work done during a summer internship.

1Johns Hopkins University 2Stanford University 3Max Planck Institute for Intelligent Systems 4MIT

Can a coding agent watch a video of a physical event and write the program that reconstructs it?

Given a video of some physical dynamics, the agent must write graphics code that reproduces that video. It has to write the geometry of the scene explicitly, its motion through time, and the rendering code that produces the final video.

200 tasks, 100 real and 100 simulated videos, 18 agents ranked

Leaderboard

How the scores are computed

    About 4DCodeBench

    4DCodeBench asks for a program rather than a prediction. Each of the 200 tasks hands an agent a single video of something physical happening: a robot folding cloth, a hydraulic press crushing a can, honey pouring, a domino line going over. It has to write code that reconstructs the event as a 4D world it can render. Half the videos are real recordings and half come from a physics simulator, so a submission can be compared against a measured scene as well as against a known one. Nothing in the task tells the agent what the objects are made of, how many there are, or where the camera sits. It has to read that off the video and commit to it in code that runs.

    Questions

    What does an agent have to produce?

    The agent writes a program and submits it together with the 4D world it generates: the rendered video, the camera, and the geometry at every frame. Running the program must regenerate everything without access to the reference video. We place no restrictions on how the motion is made; agents keyframe it, simulate it in Blender, Taichi or Warp, or write their own solvers.

    Why mix real and simulated videos?

    The two kinds of video let us evaluate different things. Real videos show realistic appearance and physical behaviour, but have no 4D ground truth, so we can only compare reconstructions of them in the image. For simulated scenes we know the exact 3D geometry and motion at every frame, so we can also compare reconstructions in 3D and over time. Both halves cover the same four families of matter: rigid and articulated bodies, deformable solids, co-dimensional structures such as cloth and rope, and flowing matter such as grains and fluids.

    What does the agent get to work with?

    Only the video and a fixed task prompt. We provide no scene metadata, object lists or descriptions, so the agent has to work out what is in the video before writing any code. Each run takes place in an isolated container with a GPU, Blender and common simulation libraries, but without the scorer. Each model gets one attempt per scene.

    89.6% of submissions run end to end. More reasoning improves results within a model: raising GPT-6 Astra’s reasoning effort from Low to Max lifts its Overall score from 0.73 to 0.79. Across different models, the number of steps or tokens used is a poor predictor of the score.

    How is a reconstruction scored?

    We compare each reconstruction with the reference video, and for simulated scenes also with the true 4D world. The metrics fall into five families, and the Overall score is their unweighted average. VQA and the two Elo ratings are reported separately.

      The Metrics tab shows each metric on a strong and a weak run of the same scene.

      What if a run fails?

      Failed runs are penalized, not omitted: a run that renders nothing or produces an invalid world receives the worst value on every metric it could not be measured on. Because models fail on different scenes, filtering out unsuccessful runs would evaluate models on mismatched subsets of the benchmark.

      How are the Elo ratings produced, and can the VLM judge be trusted?

      Both ratings come from pairwise comparisons. The judge sees the reference video and two anonymised reconstructions and picks the one that matches it better. For VLM Elo the judge is a vision-language model; for human Elo, 76 participants made 3,587 such judgments across all 200 scenes.

      The VLM’s judgments are close to human judgments. The two rankings are nearly identical (Spearman ρ = 0.975), and on individual comparisons the VLM agrees with humans almost as often as humans agree with each other (89.3% vs. 92.1%). Individual VLM judgments are close to chance for models within about 50 Elo points of each other, and become more reliable as the gap widens. The Human agreement tab plots the two ratings against each other.

      Can I evaluate my own agent?

      Yes. We release the harness, the data and the judging scripts; see Run the benchmark below. Scoring runs in a separate container, so the ground truth cannot leak into a submission.

      Run the benchmark

      A typical run consists of two main stages. First, the agent reconstructs the input video's scene into code within an isolated container, which holds nothing from the case except that video. Second, the generated output is evaluated against the reference in a separate scoring container.

      For configuration formats, agent credentials, and a complete list of options, please refer to the README.

      1. Setup

      Download the data and the checkpoints, and build the three images:

      bash
      git clone https://github.com/4DCodeBench/4DCodeBench && cd 4DCodeBench
      python scripts/download_data.py     # benchmark videos, masks and reference worlds
      scripts/download_checkpoints.sh     # the models the scorer measures with
      docker build -t 4dcb/sim-base:dev -f harness/base/Dockerfile .
      docker build -t 4dcb/agent:dev    -f harness/agent/Dockerfile .
      docker build -t 4dcb/scorer:dev   -f harness/scorer/Dockerfile .

      Without Docker, for example on a cluster, Apptainer or SingularityCE works too: convert each image to a SIF file and set backend = "sif" in the runtime file. Agents always run in a container, but scoring can also run directly in the conda environment.

      2. Run an agent

      Choose the cases, the agent and the model in the jobs file, and set paths, images and credentials in the runtime file:

      bash
      mkdir -p .local
      cp harness/runtime/templates/jobs.example.toml    .local/jobs.toml      # cases, agent, model
      cp harness/runtime/templates/runtime.example.toml .local/runtime.toml   # paths, images, credentials
      python harness/runtime/infer.py --jobs .local/jobs.toml --runtime .local/runtime.toml
      • Claude, GPT and Gemini models run through their own CLIs, signed in with a credential file in .auth/. Any other model with an OpenAI-compatible endpoint runs through Artificial Analysis' Stirrup harness, with the endpoint's URL and key set in the runtime file.
      • Each run has a wall-clock limit of six hours by default.

      3. Score

      Score what the agent delivered against the reference:

      bash
      python harness/runtime/score.py --jobs .local/jobs.toml --runtime .local/runtime.toml --stage prepare,score
      • prepare computes each case's reference estimates, once per case.
      • score computes the metrics of each run and writes them to the run's results/.

      4. VLM judge

      Two judged evaluations complement the metrics: VQA, which asks questions about one agent's renders, and a pairwise preference between two agents, fitted to an Elo rating. Both need an OpenRouter key.

      bash
      export OPENROUTER_API_KEY=...
      python vlm_judge/judge.py vqa   --system <agent>-<model>-<effort> --out answers.jsonl
      python vlm_judge/judge.py score answers.jsonl

      To rate a new agent against the agents in the paper, place their renders under runs/ first:

      bash
      hf download 4DCodeBench/Results --repo-type dataset --local-dir results
      python scripts/place_renders.py results
      python vlm_judge/judge.py pairwise --out pairs.jsonl
      python vlm_judge/judge.py elo      pairs.jsonl

      Citation

      bibtex