Chenxiang ZHANG

PhD
University of Luxembourg

Modern deep learning has shifted from MLPs and CNNs to large Transformers, together with optimizers such as SGD to Adam. These gains in performance come with models that are increasingly harder to train and to understand. This doctoral project investigates the empirical learning dynamics of such systems: How loss and gradients propagate through deep networks, and how the optimization trajectory shapes a model's final properties, along two axes: optimization and architecture.

On the optimization axis, the project characterizes how the optimizer's shapes the "global" loss landscape between different solutions (i.e. model merging). With higher noise yielding greater stability when interpolating between different models in the non-convex landscape. It further revisits the long-standing Adam-SGD performance gap and, rather than attributing it to any single factor, identifies a crossover batch size at which the relative advantage shifts from SGD to Adam, reconciling competing single-factor explanations across diverse datasets and both Transformer and non-Transformer architectures.

On the architecture axis, ongoing and planned work seeks a unified, dynamics-based account of why components such as gated linear units and normalization layers are so effective and so widely adopted. Together, these threads aim to establish the learning dynamics, rather than the optimizer or architecture taken in isolation, as a central lens for understanding the mergeability, trainability, and design of modern neural networks.

Academic Track
October 1st, 2023 - October 1st, 2027
ELLIS Edge Newsletter
Join the 6,000+ people who get the monthly newsletter filled with the latest news, jobs, events and insights from the ELLIS Network.