Aleksandr Razin
PhD
Institute for Computer Science, Artificial Intelligence and Technology (INSAIT)

Long-horizon robotic tasks are difficult because an agent needs to remember what it has observed, track how the environment changes, and notice when an action does not produce the expected result. Current vision-language-action policies are often focused on predicting the next action, so errors can build up over a longer sequence, especially when the environment is only partially observed. Generative world models could help by predicting what is likely to happen after an action, but their predictions may not always agree with the actual geometry, object states, or physical constraints of the scene.

The project would explore combining generative world prediction with a persistent 3D representation of the environment. The 3D state would keep track of task-relevant information such as objects and their poses, spatial relations, articulated parts, and interaction affordances. A generative model would predict possible outcomes of actions or subgoals, while a pretrained VLA policy would be used for low-level control. The visual agent would use this information to choose subgoals, check whether an action worked, update its state, and replan after a failure.

The main question is whether using both generative prediction and a structured 3D state makes long-horizon tasks more reliable than using either one alone. As an initial evaluation, I propose to use long-horizon household tasks from BEHAVIOR-1K and additionally test compositional generalization on RoboCasa365 Composite-Unseen.

Academic Track
November 1st, 2026 - October 31st, 2030
ELLIS Edge Newsletter
Join the 6,000+ people who get the monthly newsletter filled with the latest news, jobs, events and insights from the ELLIS Network.