Mobile visual scene understanding
RWTH Publications (RWTH Aachen)
Abstract
This thesis is concerned with the problem of mobile visual scene understanding, that is of recovering the geometry and semantics of the scene inside which a mobile platform (e.g. a vehicle or robot) navigates. This problem becomes increasingly important, as the development of autonomous driving cars and mobile robots has recently been transformed into a reality. Furthermore, there has been a substantial amount of research on simpler individual components (i.e. semantic segmentation, depth estimation, etc.), bringing them to a robustness level, which allows them to be used as a basis for building higher-level scene understanding algorithms. The scene understanding problem can be divided into two parts: the geometric reconstruction and the semantic segmentation of the scene. These two parts have been treated individually in separate lines of research, achieving impressive results in each of the sub-problems. However, the interactions and benefits that each of these problems can gain from a joint treatment have not been thoroughly explored. Moreover, this joint optimization is extremely important for mobile scenarios, where challenges demand the use of all available information sources. This thesis considers the problem of scene understanding from mobile platforms as a unique one and builds a tight interplay between the scene labeling and the geometric reconstruction components. The core part of the work is constituted of a probabilistic framework which couples the semantic labeling of consecutive video frames via the underlying 3D reconstruction. As shown in our experiments, the coupling between these two processes and the enforcement of temporal consistency in the semantic labels, allow both of the components to benefit and improve their individual performances. The resulting system creates semantic reconstructions out of a video stream captured from a mobile platform. In addition, we also explore the use of freely available street map data towards a more consistent scene representation. An important contribution in this direction is the development of a localization algorithm which registers the trajectory of a mobile platform on the street map. As our experimental results indicate, a mobile platform can be accurately localized in a street map, conferring the possibility for bidirectional information flow between the map and local reconstructions. Furthermore, we explore the use of semantic information to improve the platform's localization accuracy and in return we take advantage of the street map data to provide more detailed semantic labels. An extensive evaluation on several large datasets suggests the proposed system's applicability in real-world problems.
Authors 0
- Author list not loaded yet.
Cited by 0 stored of 0
No patents citing this paper on Lens.org (checked 2026-10-06).