Ana-Maria Marcu
PhD
University College London (UCL)

Autonomous agents, such as humanoid robots or self-driving vehicles have seen major improvements over the last decade, with public demonstrations of impressive capabilities. The emergence of the transformer architecture has enabled seamless integration of multiple modalities such as vision, language, and action into a single model. In addition, moving away from simple imitation learning methods towards reinforcement learning, human capabilities have been exceeded at games such as chess and Go. One particular feature of such games is that the state and action spaces are finite, and that the current state is sufficient for determining the optimal next action. This is, however, not true for many embodied intelligence tasks such as cooking, cleaning, or driving, where past states and actions are crucial for optimal decision-making. One limitation of existing architectures, such as transformers, linear transformers and state-space models is that they lack an explicit state update mechanism that enforces minimal sufficient statistics under partial observability. This PhD thesis aims to investigate memory mechanisms for partially observable reinforcement learning that can efficiently operate over very long visual sequences.

Industry Track
October 1st, 2026 - October 1st, 2030
ELLIS Edge Newsletter
Join the 6,000+ people who get the monthly newsletter filled with the latest news, jobs, events and insights from the ELLIS Network.