A diagnostic benchmark · visual world modeling

Can image modelsimagine time?

Image generators now paint stunning single frames — but can they hold a coherent visual world across multiple states in one picture? ImageTime asks a model to draw a single image of four ordered key states and probes whether identities, objects, spatial relations, and causal order survive over time.

1University of Electronic Science and Technology of China  ·  *Equal contribution

The paper

Read the whole paper, right here.

The full preprint — no download required. Scroll through it below, or open it full-screen / save a copy from the toolbar.

ImageTime · arXiv:2606.10620 · preprint Open full-screen ↗  ·  Download
The ImageTime Benchmark 750 cases · 22 domains · 8 models
ImageTime benchmark overview: thousands of generated motion sheets across 22 domains
At a glance prompt-only · single generation · no cherry-picking
750
benchmark cases
22
domains
375
action concepts
8
models evaluated
4 ordered key states / image L0–L6 capability tree C0–C9 capability scores D0–D14 diagnostic subscores GPT-5.5 VLM-as-judge

The question

We've long assumed image models render how things look, and video models render how things change. This generation of image models is quietly blurring that line.

Ask a strong model to draw “open a drawer → reach in → take out an object” as four stages in one image, and it doesn't just paint four pretty pictures — it keeps the object from appearing in the hand before the drawer is open. In static pixels, it maintains a world with order and causality.

But existing text-to-image benchmarks mostly score a single image: right objects, attributes, counts, layout. The moment a task involves change over time, that lens is easy to fool — a model can jump straight to the result, copy-paste a frame, conjure objects from nowhere, or let an effect precede its cause. Each frame looks fine alone; only the ordered process exposes the failure.

So instead of generating dense video and evaluating that, ImageTime forces the model to draw time into one checkable static image.

The task

One image, four ordered states.

Given an action instruction — and optionally a reference image for the initial state — a model must produce a single image laid out as a 2×2 motion sheet:

t₁ initial state t₂ action onset t₃ transition t₄ final state
The challenge — 4-keyframe spatiotemporal consistency 5 diagnosed failure modes
ImageTime task interface: a detailed prompt, four ordered key states, and five diagnosed failure modes, with the L0–L6 capability tree
premature final state
Result before cause

The outcome appears before the necessary actions ever happen.

missing interaction
No real contact

The key interaction (the hit, the grasp, the touch) is skipped.

identity / scene drift
World not preserved

The subject's face, the background, or the camera silently changes.

object duplication
Counts break

Objects appear or vanish; quantities aren't conserved across frames.

causal-order violation
Effect precedes cause

A reward shows up before the event that should produce it.

How we score

A seven-level capability tree.

Rather than one blurry total, ImageTime maps each scenario onto a progressive hierarchy where higher levels depend on the visual promises of lower ones — so you can see which level a model starts to fail.

L0
Static
Grounding the scene
L1
Identity
Anchors preserved
L2
Spatial
Relations consistent
L3
Object
State transitions
L4
Interaction
Contact & affordance
L5
Causal
Process composition
L6
Constraint
Counterfactual rules

low-level grounding → … → constrained, evolving causal worlds

C0 – C9
Capability scores

Did the model do it? Layout, entity consistency, spatial coherence, motion continuity, temporal order, causality, interaction, constraint sensitivity, and image quality.

D0 – D14
Diagnostic subscores

Why did it succeed or fail? Concrete visual evidence — is the transition visible, are counts conserved, is occlusion consistent — behind every judgement.

GPT-5.5
Structured VLM-as-judge

First validates the 2×2 layout, then parses each cell, then emits C/D scores, confidence, and failure labels — so a low score is traceable, not a black box.

Leaderboard · prompt-only

Eight models, one honest test.

No reference image, single generation, no retries or human cherry-picking. Static scenes are easy; maintaining a constrained, evolving causal world is hard.

ModelL0L1L2L3L4L5L6
GPT Image 28.117.447.457.457.567.627.28
Nano Banana 27.667.007.197.107.307.465.83
Seedream 5.0 Lite7.556.716.676.626.826.736.73
FLUX.2 Pro6.936.545.645.185.354.484.52
Qwen-Image-25126.355.944.864.634.933.601.21
Z-Image-Turbo6.356.484.964.554.553.291.44
HunyuanImage-2.15.644.544.394.034.743.791.20
SDXL2.992.131.211.371.260.490.61

Tree-level scores (L0 Static → L6 Constraint), prompt-only — the exact numbers behind the degradation curve. Note how every model except GPT Image 2 and Seedream drops sharply at L6.

Score along the capability treeL0 → L6
Tree-level degradation curve: every model's score from L0 Static to L6 Constraint, with steep drops at the harder causal and constraint levels

From static grounding (L0) to constraint reasoning (L6). The strongest models stay flat; weaker ones show clear upper-tree collapse — causality and constraints are where things break.

01

Looks ≠ understanding

Many models keep high image quality (C9) while motion continuity, temporal order and causality (C4–C6) collapse. We never let pretty pixels compensate for missing causality.

02

Not a closed-source privilege

Open models without explicit chain-of-thought (Qwen-Image, Z-Image) still show non-trivial entity preservation, spatial anchoring and ordered state change. They trail the best — but they're in the race.

03

The top of the tree is fragile

Almost every weaker model degrades upward: building a static scene is easy; sustaining a constrained, evolving causal world is hard. Failures recur — drift, missing mid-states, premature endings.

Deeper analysis

Where models break, in detail.

Beyond a single ranking, the structured judge yields capability profiles, per-diagnostic breakdowns, score distributions, cost trade-offs and domain-level maps. The headline numbers sit on the left; flip through the supporting figures in the viewer below.

ModelC meanD meanOverall
1GPT Image 27.867.877.86
2Nano Banana 27.437.477.45
3Seedream 5.0 Lite7.137.207.16
4FLUX.2 Pro5.926.286.10
5Z-Image-Turbo5.145.695.41
6Qwen-Image-25125.095.555.32
7HunyuanImage-2.14.915.044.98
8SDXL1.491.641.57

Mean capability (C), diagnostic (D) and overall scores over all 750 cases. Cell shading runs warm (strong) → cool (weak).

Capability profileC0 – C9
Radar chart of C0-C9 capability scores for all eight models

The C0–C9 capability profile. The top systems span a wide, even polygon; weaker models cave in on motion, temporal and causal axes while keeping layout and quality.

Diagnostic breakdown · D0–D14 1 / 5
Diagnostic breakdown · D0–D14

Per-diagnostic mean scores across all eight models. Everyone parses the 2×2 panel (D0) and renders legibly (D14); the real bleeding is on state transition (D8), phase ordering (D9) and constraint adherence (D13).

Side by side

Same prompt, eight motion sheets.

For one action, every model must draw its own 2×2 sheet. Strong models break the action into a coherent ordered process; weaker ones drift, repeat a stage, or collapse the layout. Below: all 22 domains, two examples each — 44 motion sheets, eight models per sheet. Filter by domain, click any image to enlarge.

What it means

Your next-generation video model — need it even be a video model?

A careful caveat first: we don't claim to see a “world model” inside these networks. ImageTime is a behavioral probe — it watches whether world-model-like consistency is externalized into the generated image, not the internal representation.

Even so: if an image model can hold a single visual world across frames, understand object state change, and respect temporal and causal order, it becomes a lightweight way to express a visual process — storyboard keyframes, robotic action previews, low-cost process-level data.

Push the frame count high enough — 30, 60 cells — flip through them, and you've essentially assembled a video out of one image model's keyframes. Between image and video generation, perhaps only “frame count” remains.

Citation

Cite ImageTime.

Our data, generated images, judge scores and scripts are fully open. Check it, reproduce it, push back on it.

BibTeX
@misc{wu2026imagemodelsimaginetime,
  title        = {Can Image Models Imagine Time? ImageTime: A Novel Benchmark
                  for Probing Visual World Modeling Through Spatiotemporal Consistency},
  author       = {Xinrui Wu and Lichen Huang},
  year         = {2026},
  eprint       = {2606.10620},
  archivePrefix= {arXiv},
  primaryClass = {cs.CV},
  url          = {https://arxiv.org/abs/2606.10620}
}