David Guzman Piedrahita

PhD
University of Zurich (UZH)

Incentives are everywhere. Wherever there are goals and limited means, they appear. Rule-following and cheating, cooperation and conflict all can emerge as agents respond to incentives originating from an environment or social interactions.

Such emergent behaviors are increasingly relevant to the safety of modern Al. The field of Al safety addresses risks by evaluating how systems behave and developing techniques to keep behavior aligned to human goals, especially in conditions model designers did not foresee. Foregrounding incentives and interactions lets us study these behaviors systematically.

First, it lets us study the robustness and dynamics of aligned values. Alignment is largely instilled and evaluated in single-agent settings, so interaction is a regime shift under which it can fail to generalize. In deployment settings that are quickly emerging yet underexplored, we can ask when aligned behavior holds and when it drifts under incentive and adversarial pressure. Second, the resulting safety findings can inform new, nonstandard alignment recipes. For example, approaches that in addition to subject matter, like bio and cyber risks, can account for incentive structures that can incentivize undesirable behavior, like miscoordination and collusion.

These threads run from the standard single-agent case, to the multi-agent case where emergent coordination is at once a safety risk and a route to more capable and controllable systems. As deployment increasingly is interaction, understanding how incentives shape emergent behavior becomes a precondition for keeping these systems aligned.

Academic Track
September 1st, 2026 - August 31st, 2031
ELLIS Edge Newsletter
Join the 6,000+ people who get the monthly newsletter filled with the latest news, jobs, events and insights from the ELLIS Network.