ECCV 2026 Workshop · Malmö

World Models in the Lp:
Towards Application-Driven World Model Evaluation

Tuesday, September 8 — Half-Day Workshop (pm)

Room: TBA

01

Overview

The next challenge for intelligent systems is not only perception but also anticipating and acting in physical space under interventions and uncertainty. World models support simulation, data generation, and planning by rolling out futures conditioned on observations and actions, in applications such as robot manipulation, autonomous driving, and embodied navigation.

Yet recent rapid progress in world models has outpaced evaluation: beyond visual fidelity, we must assess controllability, physical plausibility, and robustness to distribution shift. As world models are getting embedded in larger systems and are effectively used in-the-loop for planning or training and evaluating perception and control policies, open-loop benchmarks provide limited insight. In such settings, errors compound, and missing controllability, physical plausibility, or robustness under distribution shift directly translate into inaccurate rollouts, suboptimal decisions, degraded downstream performance, or potentially catastrophic errors in safety-critical applications. Such properties are task-dependent and, thus, require task-specific evaluation criteria.

02

Topics of Interest

Our workshop rethinks world model evaluation from an application perspective and aims to address key open questions:

01

Key capabilities under interventions

What are key capabilities of world models and their trade-offs? E.g., efficiently predicting plausible futures under interventions, generating conditioned high visual-fidelity videos, or serving as a realistic simulator for closed-loop agent training.

02

Evaluation protocols and benchmarks

How to design evaluation protocols and benchmarks for properties such as controllability, physical plausibility, and robustness? And how can exploration of these insights inform world model training and design?

03

Failure modes in downstream systems

How to identify systematic failure modes that emerge when world models are embedded in downstream systems and how to integrate imperfect world models, especially under challenging scenarios?

03

Call for Papers

We invite submissions of research papers related to the application of world models in other downstream tasks, their benchmarking, and evaluation. This is a Nectar track, i.e., papers will not be published in the ECCV 2026 Workshop Proceedings. The goal is to bring together researchers from different communities in a poster session and to promote papers in this field that were already accepted in or submitted to a previous conference (CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, RSS, CoRL).

Submissions are handled via this Google form. The submission should include a reference to the already published paper or a single PDF containing the final version (not anonymized). Submissions are processed on a rolling basis, with feedback within a week after submission.

All accepted papers will be presented in a poster session.

Important Dates:

Submission Deadline
August 14, 2026
Acceptance Decision
Rolling basis within 1 week
04

Schedule

13:30 – 13:40 Welcome / Opening Remarks
13:40 – 14:10 Invited Talk 1: Speaker 1
14:10 – 14:40 Invited Talk 2: Speaker 2
14:40 – 15:10 Invited Talk 3: Speaker 3
15:10 – 16:10 Poster Session
16:10 – 16:40 Invited Talk 4: Speaker 4
16:40 – 17:10 Invited Talk 5: Speaker 5
17:10 – 17:50 Panel Discussion
17:50 – 18:00 Ending Remarks
05

Invited Speakers

Laura Leal-Taixé

NVIDIA, TUM

Amir Bar

Imperial College London, AMI Labs

Carl Doersch

Google DeepMind

Ranjay Krishna

University of Washington, Microsoft AI

06

Organizers

Olaf Dünkel

MPI-INF

Nhi Pham

MPI-INF

Ken-Joel Simmoteit

TU Darmstadt

Thu Nguyen-Phuoc

Prometheus

Adam Kortylewski

CISPA Helmholtz Center

07

Advisory Board

Jan Peters

TU Darmstadt, DFKI

Alan Yuille

Johns Hopkins University