2 citations · 3 across the 5 of their papers we have counts for
5 papers
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
Junwen He, Yifan Wang, Lijun Wang +5
Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs wit…
Tracking with Human-Intent Reasoning
Jiawen Zhu, Zhi-Qi Cheng, Jun-Yan He +5
Advances in perception modeling have significantly improved the performance of object tracking. However, the current methods for specifying the target object in the initial frame a…
PoSynDA: Multi-Hypothesis Pose Synthesis Domain Adaptation for Robust 3D Human Pose Estimation
Hanbing Liu, Jun-Yan He, Zhi-Qi Cheng +8
Existing 3D human pose estimators face challenges in adapting to new datasets due to the lack of 2D-3D pose pairs in training sets. To overcome this issue, we propose \textit{Multi…
DAMO-StreamNet: Optimizing Streaming Perception in Autonomous Driving
Jun-Yan He, Zhi-Qi Cheng, Chenyang Li +5
Real-time perception, or streaming perception, is a crucial aspect of autonomous driving that has yet to be thoroughly explored in existing research. To address this gap, we presen…
HDFormer: High-order Directed Transformer for 3D Human Pose Estimation
Hanyuan Chen, Jun-Yan He, Wangmeng Xiang +6
Human pose estimation is a challenging task due to its structured data sequence nature. Existing methods primarily focus on pair-wise interaction of body joints, which is insuffici…