Score: 0

FutrTrack: A Camera-LiDAR Fusion Transformer for 3D Multiple Object Tracking

Published: October 22, 2025 | arXiv ID: 2510.19981v1

By: Martha Teiko Teye, Ori Maoz, Matthias Rottmann

Potential Business Impact:

Helps self-driving cars track objects better.

Business Areas:

Image Recognition Data and Analytics, Software

We propose FutrTrack, a modular camera-LiDAR multi-object tracking framework that builds on existing 3D detectors by introducing a transformer-based smoother and a fusion-driven tracker. Inspired by query-based tracking frameworks, FutrTrack employs a multimodal two-stage transformer refinement and tracking pipeline. Our fusion tracker integrates bounding boxes with multimodal bird's-eye-view (BEV) fusion features from multiple cameras and LiDAR without the need for an explicit motion model. The tracker assigns and propagates identities across frames, leveraging both geometric and semantic cues for robust re-identification under occlusion and viewpoint changes. Prior to tracking, we refine sequences of bounding boxes with a temporal smoother over a moving window to refine trajectories, reduce jitter, and improve spatial consistency. Evaluated on nuScenes and KITTI, FutrTrack demonstrates that query-based transformer tracking methods benefit significantly from multimodal sensor features compared with previous single-sensor approaches. With an aMOTA of 74.7 on the nuScenes test set, FutrTrack achieves strong performance on 3D MOT benchmarks, reducing identity switches while maintaining competitive accuracy. Our approach provides an efficient framework for improving transformer-based trackers to compete with other neural-network-based methods even with limited data and without pretraining.

Beyond Frame-wise Tracking: A Trajectory-based Paradigm for Efficient Point Cloud Tracking

CV and Pattern Recognition

Helps robots track moving things better, faster.

14 Sep 2025 0

88%

Multi-View 3D Point Tracking

CV and Pattern Recognition

Tracks moving things in 3D with few cameras.

28 Aug 2025 0

88%

FUSER: Feed-Forward MUltiview 3D Registration Transformer and SE(3)$^N$ Diffusion Refinement

CV and Pattern Recognition

Lets 3D scanners map places faster and better.

10 Dec 2025 3

View PDF Login to Bookmark

Page Count

10 pages

FutrTrack: A Camera-LiDAR Fusion Transformer for 3D Multiple Object Tracking

Helps self-driving cars track objects better.

Technical Abstract

Beyond Frame-wise Tracking: A Trajectory-based Paradigm for Efficient Point Cloud Tracking

Multi-View 3D Point Tracking

FUSER: Feed-Forward MUltiview 3D Registration Transformer and SE(3)$^N$ Diffusion Refinement