Online 3D semantic segmentation, which aims to perform real-time 3D scene reconstruction along with semantic segmentation, is an important but challenging topic. A key challenge is to strike a balance between efficiency and segmentation accuracy. There are very few deep-learning-based solutions to this problem, since the commonly used deep representations based on volumetric-grids or points do not provide efficient 3D representation and organization structure for online segmentation. Observing that on-surface supervoxels, i.e., clusters of on-surface voxels, provide a compact representation of 3D surfaces and brings efficient connectivity structure via supervoxel clustering, we explore a supervoxel-based deep learning solution for this task. To this end, we contribute a novel convolution operation (SVConv) directly on supervoxels. SVConv can efficiently fuse the multi-view 2D features and 3D features projected on supervoxels during the online 3D reconstruction, and leads to an effective supervoxel-based convolutional neural network, termed as
Supervoxel-CNN
, enabling 2D-3D joint learning for 3D semantic prediction. With the
Supervoxel-CNN
, we propose a
clustering-then-prediction
online 3D semantic segmentation approach. The extensive evaluations on the public 3D indoor scene datasets show that our approach significantly outperforms the existing online semantic segmentation systems in terms of efficiency or accuracy.
[1]
Ke Xie,et al.
A search-classify approach for cluttered indoor scene understanding
,
2012,
ACM Trans. Graph..
[2]
Yiwei Jin,et al.
3D reconstruction using deep learning: a survey
,
2020,
Commun. Inf. Syst..
[3]
Matthias Nießner,et al.
BundleFusion
,
2016,
TOGS.
[4]
Mingmin Zhen,et al.
Multi-view based neural network for semantic segmentation on 3D scenes
,
2019,
Science China Information Sciences.
[5]
Matthias Nießner,et al.
Real-time 3D reconstruction at scale using voxel hashing
,
2013,
ACM Trans. Graph..
[6]
Qinping Zhao,et al.
Semantic part segmentation of single-view point cloud
,
2020,
Science China Information Sciences.
[7]
Shi-Min Hu,et al.
Real-time High-accuracy Three-Dimensional Reconstruction with Consumer RGB-D Cameras
,
2018,
ACM Trans. Graph..
[8]
Daniel Cohen-Or,et al.
Semantic object reconstruction via casual handheld scanning
,
2018,
ACM Trans. Graph..
[9]
Hao Su,et al.
3D attention-driven depth acquisition for object identification
,
2016,
ACM Trans. Graph..