Agentic systems built on large language and multimodal models can plan, write code, and invoke external tools, yet their reliability degrades precisely where autonomy matters most: selecting among many capabilities, grounding answers in modalities the base model handles poorly, and executing multi-step interactions without verifiable supervision. This PhD investigates how to measure and improve these abilities, focusing on agents that must coordinate with external models and tools rather than answer from parametric knowledge alone.
This project investigates questions about what agents must be given versus what they can establish themselves: how exhaustively capabilities and constraints must be described before an agent can act competently, and whether it can discover them through interaction; how it should manage memory over long horizons; and whether it can improve continuously from its own experience.