Ege Erdogan

PhD
University of Amsterdam (UvA)
Mechanistic Interpretability for Science

The PhD is primarily on developing and applying mechanistic interpretability tools on ML models used in scientific domains such as weather prediction. Such models tend to outperform traditional non-ML methods in various ways, yet remain black boxes as their internals are not interpretable. Being able to better interpret their internal computations would help establish trust in these models, and help discover scientific insights. Mechanistic interpretability in particular aims to explain ML models' computations through their internal activations, e.g. by decomposing them into features that then build up to circuits that causally explain certain behaviors of these models.

Mechanistic interpretability methods have been widely used on language models and have seen some recent adoption for scientific models. This PhD will aim to further this advancement by 1) developing mechanistic interpretability methods better suited to the scientific problems (e.g. the temporal dynamics of weather) and 2) using those methods to interpret the computations frontier models on specific tasks such as tropical cyclone modeling.

Academic Track
May 1st, 2025 - May 1st, 2029
ELLIS Edge Newsletter
Join the 6,000+ people who get the monthly newsletter filled with the latest news, jobs, events and insights from the ELLIS Network.