4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes
∗Equal contribution, listed in random order. ‡Equal co-advising. †Work done during a summer internship.
1
2
3
4
Paper Citation Leaderboard Data Code
Can a coding agent watch a video of a physical event and write the program that reconstructs it?
Given a video of some physical dynamics, the agent must write graphics code that reproduces that video. It has to write the geometry of the scene explicitly, its motion through time, and the rendering code that produces the final video.
Leaderboard
How the scores are computed
How to read this plot
- The axesAny score against what a model uses per task: the mean over its 200 runs of the tokens it read and wrote, or of what they cost, on a log scale. Open-weight models are priced at OpenRouter rates.
- The lineThe Pareto frontier: no model uses less and scores more than one on it.
- The barsOn an Elo score, the 95% bootstrap interval of the rating.
The ratings as a table
What each column measures
About 4DCodeBench
4DCodeBench asks for a program rather than a prediction. Each of the 200 tasks hands an agent a single video of something physical happening: a robot folding cloth, a hydraulic press crushing a can, honey pouring, a domino line going over. It has to write code that reconstructs the event as a 4D world it can render. Half the videos are real recordings and half come from a physics simulator, so a submission can be compared against a measured scene as well as against a known one. Nothing in the task tells the agent what the objects are made of, how many there are, or where the camera sits. It has to read that off the video and commit to it in code that runs.
Questions
What does an agent have to produce?
The agent writes a program and submits it together with the 4D world it generates: the rendered video, the camera, and the geometry at every frame. Running the program must regenerate everything without access to the reference video. We place no restrictions on how the motion is made; agents keyframe it, simulate it in Blender, Taichi or Warp, or write their own solvers.
Why mix real and simulated videos?
The two kinds of video let us evaluate different things. Real videos show realistic appearance and physical behaviour, but have no 4D ground truth, so we can only compare reconstructions of them in the image. For simulated scenes we know the exact 3D geometry and motion at every frame, so we can also compare reconstructions in 3D and over time. Both halves cover the same four families of matter: rigid and articulated bodies, deformable solids, co-dimensional structures such as cloth and rope, and flowing matter such as grains and fluids.
What does the agent get to work with?
Only the video and a fixed task prompt. We provide no scene metadata, object lists or descriptions, so the agent has to work out what is in the video before writing any code. Each run takes place in an isolated container with a GPU, Blender and common simulation libraries, but without the scorer. Each model gets one attempt per scene.
89.6% of submissions run end to end. More reasoning improves results within a model: raising GPT-6 Astra’s reasoning effort from Low to Max lifts its Overall score from 0.73 to 0.79. Across different models, the number of steps or tokens used is a poor predictor of the score.
How is a reconstruction scored?
We compare each reconstruction with the reference video, and for simulated scenes also with the true 4D world. The metrics fall into five families, and the Overall score is their unweighted average. VQA and the two Elo ratings are reported separately.
The Metrics tab shows each metric on a strong and a weak run of the same scene.
What if a run fails?
Failed runs are penalized, not omitted: a run that renders nothing or produces an invalid world receives the worst value on every metric it could not be measured on. Because models fail on different scenes, filtering out unsuccessful runs would evaluate models on mismatched subsets of the benchmark.
How are the Elo ratings produced, and can the VLM judge be trusted?
Both ratings come from pairwise comparisons. The judge sees the reference video and two anonymised reconstructions and picks the one that matches it better. For VLM Elo the judge is a vision-language model; for human Elo, 76 participants made 3,587 such judgments across all 200 scenes.
The VLM’s judgments are close to human judgments. The two rankings are nearly identical (Spearman ρ = 0.975), and on individual comparisons the VLM agrees with humans almost as often as humans agree with each other (89.3% vs. 92.1%). Individual VLM judgments are close to chance for models within about 50 Elo points of each other, and become more reliable as the gap widens. The Human agreement tab plots the two ratings against each other.
Can I evaluate my own agent?
Yes. We release the harness, the data and the judging scripts; see Run the benchmark below. Scoring runs in a separate container, so the ground truth cannot leak into a submission.
Input video
what the agent seesAgent reconstructions
Drag across any render to compareRun the benchmark
A typical run consists of two main stages. First, the agent reconstructs the input video's scene into code within an isolated container, which holds nothing from the case except that video. Second, the generated output is evaluated against the reference in a separate scoring container.
For configuration formats, agent credentials, and a complete list of options, please refer to the README.
1. Setup
Download the data and the checkpoints, and build the three images:
git clone https://github.com/4DCodeBench/4DCodeBench && cd 4DCodeBench
python scripts/download_data.py # benchmark videos, masks and reference worlds
scripts/download_checkpoints.sh # the models the scorer measures with
docker build -t 4dcb/sim-base:dev -f harness/base/Dockerfile .
docker build -t 4dcb/agent:dev -f harness/agent/Dockerfile .
docker build -t 4dcb/scorer:dev -f harness/scorer/Dockerfile .
Without Docker, for example on a cluster, Apptainer or SingularityCE
works too: convert each image to a SIF file and set backend = "sif" in
the runtime file. Agents always run in a container, but scoring can also run directly
in the conda environment.
2. Run an agent
Choose the cases, the agent and the model in the jobs file, and set paths, images and credentials in the runtime file:
mkdir -p .local
cp harness/runtime/templates/jobs.example.toml .local/jobs.toml # cases, agent, model
cp harness/runtime/templates/runtime.example.toml .local/runtime.toml # paths, images, credentials
python harness/runtime/infer.py --jobs .local/jobs.toml --runtime .local/runtime.toml
- Claude, GPT and Gemini models run through their own CLIs, signed in with a credential
file in
.auth/. Any other model with an OpenAI-compatible endpoint runs through Artificial Analysis' Stirrup harness, with the endpoint's URL and key set in the runtime file. - Each run has a wall-clock limit of six hours by default.
3. Score
Score what the agent delivered against the reference:
python harness/runtime/score.py --jobs .local/jobs.toml --runtime .local/runtime.toml --stage prepare,score
preparecomputes each case's reference estimates, once per case.scorecomputes the metrics of each run and writes them to the run'sresults/.
4. VLM judge
Two judged evaluations complement the metrics: VQA, which asks questions about one agent's renders, and a pairwise preference between two agents, fitted to an Elo rating. Both need an OpenRouter key.
export OPENROUTER_API_KEY=...
python vlm_judge/judge.py vqa --system <agent>-<model>-<effort> --out answers.jsonl
python vlm_judge/judge.py score answers.jsonl
To rate a new agent against the agents in the paper, place their renders under
runs/ first:
hf download 4DCodeBench/Results --repo-type dataset --local-dir results
python scripts/place_renders.py results
python vlm_judge/judge.py pairwise --out pairs.jsonl
python vlm_judge/judge.py elo pairs.jsonl
Citation