Score: 2

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion

Published: August 5, 2025 | arXiv ID: 2508.03252v1

By: Wentao Qu , Guofeng Mei , Jing Wang and more

Potential Business Impact:

Finds objects in 3D scenes faster.

Denoising Diffusion Probabilistic Models (DDPMs) have shown success in robust 3D object detection tasks. Existing methods often rely on the score matching from 3D boxes or pre-trained diffusion priors. However, they typically require multi-step iterations in inference, which limits efficiency. To address this, we propose a \textbf{R}obust single-stage fully \textbf{S}parse 3D object \textbf{D}etection \textbf{Net}work with a Detachable Latent Framework (DLF) of DDPMs, named RSDNet. Specifically, RSDNet learns the denoising process in latent feature spaces through lightweight denoising networks like multi-level denoising autoencoders (DAEs). This enables RSDNet to effectively understand scene distributions under multi-level perturbations, achieving robust and reliable detection. Meanwhile, we reformulate the noising and denoising mechanisms of DDPMs, enabling DLF to construct multi-type and multi-level noise samples and targets, enhancing RSDNet robustness to multiple perturbations. Furthermore, a semantic-geometric conditional guidance is introduced to perceive the object boundaries and shapes, alleviating the center feature missing problem in sparse representations, enabling RSDNet to perform in a fully sparse detection pipeline. Moreover, the detachable denoising network design of DLF enables RSDNet to perform single-step detection in inference, further enhancing detection efficiency. Extensive experiments on public benchmarks show that RSDNet can outperform existing methods, achieving state-of-the-art detection.

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion

CV and Pattern Recognition

Finds objects in 3D space faster and better.

5 Aug 2025 2

87%

Diffusion As Self-Distillation: End-to-End Latent Diffusion In One Model

CV and Pattern Recognition

Makes AI create pictures faster and better.

18 Nov 2025 0

87%

LiDAR Point Cloud Image-based Generation Using Denoising Diffusion Probabilistic Models

CV and Pattern Recognition

Makes self-driving cars see better in bad weather.

23 Sep 2025 1

View PDF Login to Bookmark

Country of Origin

🇨🇳 China

Repos / Data Links

github.com

Page Count

18 pages

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion

Finds objects in 3D scenes faster.

Technical Abstract

Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion

Diffusion As Self-Distillation: End-to-End Latent Diffusion In One Model

LiDAR Point Cloud Image-based Generation Using Denoising Diffusion Probabilistic Models