Modern deep learning has shifted from MLPs and CNNs to large Transformers, together with optimizers such as SGD to Adam. These gains in performance come with models that are increasingly harder to train and to understand. This doctoral project investigates the empirical learning dynamics of such systems: How loss and gradients propagate through deep networks, and how the optimization trajectory shapes a model's final properties, along two axes: optimization and architecture.
On the optimization axis, the project characterizes how the optimizer's shapes the "global" loss landscape between different solutions (i.e. model merging). With higher noise yielding greater stability when interpolating between different models in the non-convex landscape. It further revisits the long-standing Adam-SGD performance gap and, rather than attributing it to any single factor, identifies a crossover batch size at which the relative advantage shifts from SGD to Adam, reconciling competing single-factor explanations across diverse datasets and both Transformer and non-Transformer architectures.
On the architecture axis, ongoing and planned work seeks a unified, dynamics-based account of why components such as gated linear units and normalization layers are so effective and so widely adopted. Together, these threads aim to establish the learning dynamics, rather than the optimizer or architecture taken in isolation, as a central lens for understanding the mergeability, trainability, and design of modern neural networks.