Autonomous systems rely on multimodal foundation models for perception, reasoning, and planning. They must interpret inputs from multiple modalities and make decisions in dynamic virtual or physical environments. However, neural networks, which underpin foundation models, are known to be highly sensitive to adversarial perturbations and distribution shifts, while their internal decision-making processes remain largely opaque. As a result, autonomous systems can become unreliable when faced with malicious attacks or rare inputs, and their behavior is difficult for humans to understand and trust. The lack of robustness and interpretability therefore constitutes a major obstacle to their trustworthy deployment.
This project investigates the foundations of trustworthy multimodal systems by jointly studying robustness, interpretability, and human understanding. These properties are fundamentally connected: robust features are more semantically aligned, reliable explanations require stable representations, and interpretability tools may enable the detection of vulnerabilities during deployment. A particular focus will be on perturbations affecting multiple modalities simultaneously and distributed over time in multi-turn interactions, which introduce novel vulnerabilities beyond traditional unimodal threat models. In addition to developing new methods to evaluate and improve robustness and interpretability, the project will investigate how humans perceive, understand, and appropriately rely on the explanations and decisions produced by multimodal systems through user-centered evaluation. By combining algorithmic advances with human-centered assessment, the project will provide a deeper understanding of the limitations of multimodal foundation models and take a significant step toward their safe, transparent, and trustworthy deployment.