Unpacking Softmax: How Logits Norm Drives Representation Collapse, Compression and Generalization

Wojciech Masarczyk, Mateusz Ostaszewski, Tin Sum Cheng, Tomasz Trzcinski, Aurelien Lucchi, Razvan Pascanu

Author Locations

No location data available for the ELLIS authors of this paper.

ELLIS Edge Newsletter
Join the 6,000+ people who get the monthly newsletter filled with the latest news, jobs, events and insights from the ELLIS Network.