This project investigates how foundation models can be extended with modular components - adapters, retrieved modules, compressed contexts, or tool/API calls - while remaining reliable, controllable, and safe. Currently, in many systems, such components are integrated without sufficient understanding about when they should actually be used in a trustworthy way. Thus, the main question is how a model, given a context, can decide whether to use a given component, abstain from using it, or fall back to a safer alternative. The initial focus may be on adapters as a controlled starting point, since they make it possible to study failure modes such as negative transfer, unexpected unsafe behavior, or poor generalization across languages and domains. More broadly, the project will explore how to identify modules that cause negative transfer or unsafe behavior, how to reject them on a per-input basis, and whether these failures are concentrated in long-tail settings such as low-resource languages, domain shifts, or personalized use cases. At a later stage, the same framework can be extended to tool-calling in agentic settings, where the main question shifts from a component being safe in general to being safe to use in the current situation.