Qi Zhang

PhD
University of Amsterdam (UvA)

Visual SLAM enables the regression of scene geometry and the generation of fine-grained semantics for text-based interaction while maintaining lightweight sensor deployment. My research focuses on enhancing semantic understanding in SLAM and maintaining system robustness against real-world camera degradation and dynamic object interference. Previous methods often rely on idealized modeling of camera degradation and tend to fail when dynamic objects occupy a large portion of the field of view. In contrast, neural network-based approaches allow for non-linear modeling and the joint estimation of both dynamic and static objects using implicit representations. My objective is to achieve robust spatial perception and semantic understanding in real-world environments by either integrating neural networks with traditional SLAM systems or fully replacing them. This will enable the system to navigate diverse and complex real-world scenarios through a comprehensive understanding of semantics and geometry.

The latest generation of SLAM systems can be categorized into three paradigms. The first utilizes monocular neural network priors combined with traditional multi-view geometric optimization. The second performs global consistency optimization on offline or stream-reconstructed scene states via geometry. The third replaces the entire pipeline with neural networks. Within this third paradigm, methods based on attention KV-cache implicit maps can jointly optimize semantic and geometric information. These methods demonstrate superior robustness over traditional approaches in certain degraded scenarios, showing great potential as the foundation for next-generation SLAM. Simultaneously, explicit point cloud storage equipped with high-dimensional semantic vectors remains a competitive representation due to its strong editability. However, current systems still struggle to scale to larger environments and are largely confined to static scenes. Existing works predominantly use the uncertainty of neural networks to filter out dynamic objects, relying solely on static cues for scene reasoning. To address this, my project proposes that performing non-linear modeling for the implicit joint estimation of both dynamic and static elements will allow the system to converge toward a more optimal and robust solution.

Academic Track
February 12th, 2024 - February 12th, 2028
ELLIS Edge Newsletter
Join the 6,000+ people who get the monthly newsletter filled with the latest news, jobs, events and insights from the ELLIS Network.