activity
20222024
most citedReading Relevant Feature from Global Representation Memory for Visual Object Tracking

5 citations · 15 across the 13 of their papers we have counts for

collaborators

13 papers

cs.CV2024

All rivers run into the sea: Unified Modality Brain-like Emotional Central Mechanism

Xinji Mai, Junxiong Lin, Haoran Wang +10

In the field of affective computing, fully leveraging information from a variety of sensory modalities is essential for the comprehensive understanding and processing of human emot…

cs.LG20243 cited

From Efficient Multimodal Models to World Models: A Survey

Xinji Mai, Zeng Tao, Junxiong Lin +5

Multimodal Large Models (MLMs) are becoming a significant research focus, combining powerful large language models with multimodal learning to perform complex tasks across differen…

cs.CV20243 cited

OneTracker: Unifying Visual Object Tracking with Foundation Models and Efficient Tuning

Lingyi Hong, Shilin Yan, Renrui Zhang +8

Visual object tracking aims to localize the target object of each frame based on its initial appearance in the first frame. Depending on the input modility, tracking tasks can be d…

cs.CV2024

OneVOS: Unifying Video Object Segmentation with All-in-One Transformer Framework

Wanyun Li, Pinxue Guo, Xinyu Zhou +5

Contemporary Video Object Segmentation (VOS) approaches typically consist stages of feature extraction, matching, memory management, and multiple objects aggregation. Recent advanc…

cs.CV2024

ClickVOS: Click Video Object Segmentation

Pinxue Guo, Lingyi Hong, Xinyu Zhou +7

Video Object Segmentation (VOS) task aims to segment objects in videos. However, previous settings either require time-consuming manual masks of target objects at the first frame d…

cs.CV2024

Align-DFER: Pioneering Comprehensive Dynamic Affective Alignment for Dynamic Facial Expression Recognition with CLIP

Zeng Tao, Yan Wang, Junxiong Lin +9

The performance of CLIP in dynamic facial expression recognition (DFER) task doesn't yield exceptional results as observed in other CLIP-based classification tasks. While CLIP's pr…