Postdoc Position for E4RL: Egocentric 4D Representations for Robot Learning at Bocconi University

Apply by 23/11/26
Career Stage: Postdoc
Milan

Today, robots are usually trained either in manually designed simulations or through expensive robot collected experience. Human videos offer a powerful alternative. They are easy to collect, widely available, and show people performing many everyday skills that we would like robots to learn. However, videos are not directly useful for robot learning. A video is only a sequence of 2D images. It shows what happened, but it does not explicitly tell the robot how the person interacted with the physical environment: where relevant objects and scene elements are in 3D, how they move, how they are used, or which of them matter for completing a task.

E4RL will address this gap by converting first-person videos into egocentric 4D representations for robot learning. Here, “first-person/egocentric” means that the video is recorded from the point of view of a person performing the task, for example using a wearable camera. “4D” means that the system builds a 3D representation of the physical scene and tracks how it changes over time. In other words, E4RL will turn videos of human activity into dynamic 3D descriptions of the physical world that robots can use for learning.

Specifically, E4RL will recover three complementary elements from first-person video: (i) the scene structure, (ii) the objects involved in the activity, and (iii) the pose and motion of the person. The goal is to recover not only what is present, but where scene elements and objects are, how they move, how they are used, and how they contribute to the task. These representations will be simple, interpretable, and efficient. By capturing the scene, the objects, and the human motion, they will provide robots with a structured description of how tasks are performed in the physical world.

The final stage of the project will utilize those scene-object-human representations as training data for robotics policies and finally deploy the learned policies on a physical humanoid robot using only its on-board sensing and computation.

This is an exciting opportunity to work at the frontier of egocentric video understanding, 3D scene representation, 4D scene understanding, and physical robot learning. The project will also benefit from the expertise of Prof. Francis Engelmann (Università della Svizzera italiana, link to bio: https://francisengelmann.github.io/) on compact 3D scene and object representations, and of Prof. Antonio Loquercio (University of Pennsylvania, link to bio: https://antonilo.github.io/) on robot policy learning and humanoid deployment.

ELLIS Edge Newsletter
Join the 6,000+ people who get the monthly newsletter filled with the latest news, jobs, events and insights from the ELLIS Network.