13 citations · 26 across the 9 of their papers we have counts for
7 papers · 1 filter
Decoupling Vision and Language: Codebook Anchored Visual Adaptation
Jason Wu, Tianchen Zhao, Chang Liu +7
Large Vision-Language Models (LVLMs) use their vision encoders to translate images into representations for downstream reasoning, but the encoders often underperform in domain-spec…
Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
Jing Tan, Zhaoyang Zhang, Yantao Shen +6
We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects…
Hyperbolic Learning with Synthetic Captions for Open-World Detection
Fanjie Kong, Yanbei Chen, Jiarui Cai +1
Open-world detection poses significant challenges, as it requires the detection of any object using either object class labels or free-form texts. Existing related works often use…
MeMOT: Multi-Object Tracking with Memory
Jiarui Cai, Mingze Xu, Wei Li +4
We propose an online tracking algorithm that performs the object detection and data association under a common framework, capable of linking objects after a long time span. This is…
ACE: Ally Complementary Experts for Solving Long-Tailed Recognition in One-Shot
Jiarui Cai, Yizhou Wang, Jenq-Neng Hwang
One-stage long-tailed recognition methods improve the overall performance in a "seesaw" manner, i.e., either sacrifice the head's accuracy for better tail classification or elevate…
Multi-Target Multi-Camera Tracking of Vehicles using Metadata-Aided Re-ID and Trajectory-Based Camera Link Model
Hung-Min Hsu, Jiarui Cai, Yizhou Wang +2
In this paper, we propose a novel framework for multi-target multi-camera tracking (MTMCT) of vehicles based on metadata-aided re-identification (MA-ReID) and the trajectory-based…