MS-Occ: Multi-Stage LiDAR-Camera Fusion for 3D Semantic Occupancy Prediction
Wei, Zhiqiang, Zheng, Lianqing, Liu, Jianan, Huang, Tao, Han, Qing Long, Zhang, Wenwen, and Zhang, Fengdeng (2026) MS-Occ: Multi-Stage LiDAR-Camera Fusion for 3D Semantic Occupancy Prediction. IEEE Robotics and Automation Letters, 11 (1). pp. 370-377.
|
PDF (Published Version)
- Published Version
Restricted to Repository staff only |
Abstract
Accurate 3D semantic occupancy perception is essential for autonomous driving in complex environments with diverse and irregular objects. While vision-centric methods suffer from geometric inaccuracies, LiDAR-based approaches often lack rich semantic information. To address these limitations, MS-Occ, a novel multi-stage LiDAR-camera fusion framework which includes middle-stage fusion and late-stage fusion, is proposed, integrating LiDAR's geometric fidelity with camera-based semantic richness via hierarchical cross-modal fusion. The framework introduces innovations at two critical stages: (1) In the middle-stage feature fusion, the Gaussian-Geo module leverages Gaussian kernel rendering on sparse LiDAR depth maps to enhance 2D image features with dense geometric priors, and the Semantic-Aware module enriches LiDAR voxels with semantic context via deformable cross-attention; (2) In the late-stage voxel fusion, the Adaptive Fusion (AF) module dynamically balances voxel features across modalities, while the High Classification Confidence Voxel Fusion (HCCVF) module resolves semantic inconsistencies using self-attention-based refinement. Experiments on two large-scale benchmarks demonstrate state-of-the-art performance. On nuScenes-OpenOccupancy, MS-Occ achieves an Intersection over Union (IoU) of 32.1% and a mean IoU (mIoU) of 25.3%, surpassing the state-of-the-art by +0.7% IoU and +2.4% mIoU. Furthermore, on the SemanticKITTI benchmark, our method achieves a new state-of-the-art mIoU of 24.08%, robustly validating its generalization capabilities. Ablation studies further confirm the effectiveness of each individual module, highlighting substantial improvements in the perception of small objects and reinforcing the practical value of MS-Occ for safety-critical autonomous driving scenarios.
| Item ID: | 91794 |
|---|---|
| Item Type: | Article (Refereed Research - C1) |
| ISSN: | 2377-3766 |
| Keywords: | 3D semantic occupancy, autonomous driving, deformable attention, LiDAR-camera fusion, multi-stage fusion, sensor fusion |
| Copyright Information: | © 2025 IEEE. All rights reserved, including rights for text and data mining, and training of artificial intelligence and similar technologies. |
| Date Deposited: | 24 Aug 2026 03:49 |
| FoR Codes: | 46 INFORMATION AND COMPUTING SCIENCES > 4603 Computer vision and multimedia computation > 460303 Computational imaging @ 100% |
| SEO Codes: | 28 EXPANDING KNOWLEDGE > 2801 Expanding knowledge > 280110 Expanding knowledge in engineering @ 100% |
| Downloads: |
Total: 2 Last 12 Months: 2 |
| More Statistics |
