Abstract
How can robots learn to understand and predict the future 3D world from vision alone? World models are emerging as a key paradigm in robotics, enabling intelligent systems to reason about how their surroundings evolve over time. This capability is particularly important for dynamic robotic applications, such as autonomous vehicles, where anticipating future scene geometry, semantics, and motion is essential for safe interaction and decision-making. Semantic 4D occupancy provides a structured representation of this evolving 3D world by jointly modeling scene geometry, semantics, and dynamics. A major challenge, however, is supervision: current occupancy models often rely on expensive, densely annotated 3D voxel data, which is difficult to obtain and scale.
In this Master's thesis, you will investigate how vision foundation models such as DINOv2, CLIP, or SAM can act as scalable semantic teachers for a vision-centric occupancy world model. By transferring rich semantic knowledge from pre-trained 2D models into spatio-temporal 3D/4D representations, the goal is to reduce the reliance on dense 3D annotations while maintaining accurate future scene prediction.
You will develop and evaluate a foundation-model-guided predictive occupancy world model using multi-view camera sequences. The thesis will explore semantic feature alignment and knowledge distillation, with the goal of enabling scalable 4D occupancy forecasting for dynamic robotic environments, such as autonomous driving.
These tasks interest you
- Develop a vision-centric occupancy world model based on Transformer architectures for predicting future semantic 4D occupancy from sequential multi-view camera inputs.
- Build and train the PyTorch pipeline, designing alignment mechanisms to distill semantic features from 2D foundation models into your 4D spatio-temporal world representation.
- Benchmark against fully-supervised baselines on large-scale datasets (e.g., nuScenes), focusing on forecasting accuracy (IoU), semantic precision, and label efficiency.
That makes you stand out
- You are registered in a master's program in computer science, artificial intelligence, robotics, or a related field.
- You have excellent programming skills in Python as well as solid experience with deep learning frameworks (especially PyTorch).
- You have a solid background in 3D computer vision. Practical experience with semantic segmentation, occupancy networks, or 3D Gaussian splatting is a major plus.
- You have knowledge of Vision Transformers (ViT), Foundation Models (DINO, CLIP), and paradigms of self- and weakly-supervised learning.
- You work independently and are solution-oriented, highly motivated, and have very good German and English skills (at least C1 level) to ensure clear and confident communication within the team and with our partners.
Your contact person
Daniela
+49 821 885882-0 work@xitaso.com