Journal article

Multi-granularity 3D Visual Language Tracking from monocular videos: A benchmark and model

Hongkai Wei, Rong Wang, Keyu Guo, Yongle Huang, Shijie Sun, Xiangyu Song, Mingtao Feng, Naveed Akhtar

Pattern Recognition | Elsevier BV | Published : 2027

Abstract

Object tracking, a fundamental task in computer vision, has recently evolved beyond purely visual cues to incorporate natural language as high-level semantic guidance. Monocular 3D Visual Language Tracking (Mono3DVLT) extends this paradigm into the 3D domain, seeking to localize and track a referred 3D object in monocular video under natural language guidance. Despite its potential, existing approaches remain constrained by single-granularity language descriptions and weak temporal reasoning. Motivated by the diversity of natural instructions in real-world traffic scenes, we introduce a comprehensive framework for evaluating whether monocular 3D trackers perform under controlled variations i..

View full abstract

University of Melbourne Researchers