Carlo Fabrizio
The real world is intrinsically non-stationary while most Machine Learning (ML) models and frameworks are designed for stationary environments. This discrepancy is even more evident in Reinforcement Learning (RL) systems, in which, even if the environment is stationary and fully observable, the agent's policy continuously changes, actively drifting the data distribution of the agent's replay buffer. The usual ML pipeline encompasses two very distinct phases, which are training and deployment. After the deployment phase, the model typically remains untouched. The naive re-adaptation process of an already deployed ML system to a new environment dynamics consists in retraining everything from scratch, which becomes unfeasible in modern machine learning systems characterized by expensive training phases. This raises the need for novel ML pipelines that enable Lifelong Learning (LL) systems. Ideally, Artificial Intelligence (AI) systems should be able to continually learn for an unlimited horizon, adapting autonomously to new operating regimes while keeping expertise in past ones. Memory is undoubtedly an essential component for such Continual Learning (CL) paradigm, and can be implemented in several ways, either explicitly through an experience memory buffer or implicitly within the learning process (e.g. regularization). This PhD will focus on the design of novel memory components for CL systems, with a particular emphasis on Continual Reinforcement Learning (CRL). The research will be balanced between theoretical and practical work, with applications to digital twins of power grids and urban traffic systems within the context of the WEL-T project (https://welri.org/cms/c_16940984/en/wel-t), whose continuoisly evolving environments make them a natural application domain for Continual Reinforcement Learning.