Multimodal large language models (MLLMs) achieve strong performance on many vision-language tasks, yet they continue to exhibit systematic failures in fine-grained perception and reasoning that remain poorly understood. The PhD thesis investigates where and why these failures occur, and develops methods to diagnose and address them across the perception-to-reasoning pipeline.
The first part introduces fine-grained, localized vision-language representations and a cross-modality self-distillation pretraining objective that improve the granularity of visual understanding available to downstream MLLMs. Building on these representations, the thesis diagnoses specific failure modes: a training-free uncertainty-guided method exposes fine-grained perceptual weaknesses without additional supervision, a hallucination benchmark reveals systematic errors under fine-grained negative queries, and a layerwise mechanistic analysis localizes where segmentation and spatial capacity emerge and degrade within the model. Moving from diagnosis toward reasoning, the thesis shows that frozen MLLMs contain latent, task-dependent computation that standard inference fails to exploit, that inference-time layer recursion recovers this capacity, and that reinforcement learning implicitly learns a related recursion policy. Finally, the thesis turns to spatial reasoning, introducing a diagnostic benchmark that disentangles genuine 3D spatial understanding from language-level shortcuts, and proposing methods that ground MLLMs in explicit 3D geometry.
Together, these contributions trace a coherent path from fine-grained perception, through systematic diagnosis, to efficient and spatially grounded reasoning in multimodal large language models.