5 citations · 6 across the 3 of their papers we have counts for
3 papers
cs.CV2024
Multi-modal Instruction Tuned LLMs with Fine-grained Visual Perception
Junwen He, Yifan Wang, Lijun Wang +5
Multimodal Large Language Model (MLLMs) leverages Large Language Models as a cognitive framework for diverse visual-language tasks. Recent efforts have been made to equip MLLMs wit…
cs.CV2023★ 1 cited
Towards Deeply Unified Depth-aware Panoptic Segmentation with Bi-directional Guidance Learning
Junwen He, Yifan Wang, Lijun Wang +6
Depth-aware panoptic segmentation is an emerging topic in computer vision which combines semantic and geometric understanding for more robust scene interpretation. Recent works pur…
cs.CV2023★ 5 cited
Tracking Anything in High Quality
Jiawen Zhu, Zhenyu Chen, Zeqi Hao +9
Visual object tracking is a fundamental video task in computer vision. Recently, the notably increasing power of perception algorithms allows the unification of single/multiobject…