VGCRTrack: Multi-Camera 3D Tracking with View-Aware Geometric Center Refinement
T B Phan, Duc-Duy Dinh, Thanh-Phong Huynh, Quoc-Thinh Le, Huy-Hoang Dang, Vu-Hoang Tran, Van-Tin Luu, Ching-Chun Huang
- Year
- 2025
- Citations
- 2
Abstract
Multi-camera multi-object tracking is a critical capability for intelligent surveillance and 3D scene understanding, enabling consistent identity tracking across disjoint views in real time. However, challenges such as severe occlusion, class ambiguity (e.g., humans vs. humanoid robots), and viewpoint variation hinder robust association-especially in online settings with limited temporal context. To address these challenges, we propose an online multi-camera 3D multiclass tracking framework. Our framework estimates 3D bounding boxes and yaw orientations by leveraging depth maps, 2D detections, and calibrated camera geometry. Orientation is inferred via a hybrid strategy combining keypoint-based estimation and a lightweight model. For robust cross-view association, we introduce two novel affinity measures: (1) a trajectory-level Fréchet distance for temporal consistency, and (2) a view-wise 3D IoU for spatial alignment. To further improve localization, we propose the View-Aware Geometric Center Refinement (VGCR) module that fuses multi-view depth and temporal cues to refine 3D center estimation. Our method achieves a HOTA score of 25.3983% on Track 1 of the 2025 AI City Challenge [9], ranking 4th on the public leaderboard.
Keywords
Related papers
Statistical Learning Theory
Yuhai Wu, Vladimir Vapnik
1999
Fractional Differential Equations
Igor Podlubný
2025
Applied Nonlinear Control
Jean-Jacques Slotine, Weiping Li
1991
Genetic Programming: On the Programming of Computers by Means of Natural Selection
John R. Koza
1992