9 papers
ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models
Longyu Chen, Heng Li, Wei Yang +2
Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for advancing embodied intelligence, enabling the safety-critical evaluation of Vision-Language-Actio…
ReGLA: Efficient Receptive-Field Modeling with Gated Linear Attention Network
Junzhou Li, Manqi Zhao, Yilin Gao +4
Balancing accuracy and latency on high-resolution images is a critical challenge for lightweight models, particularly for Transformer-based architectures that often suffer from exc…
SaFiRe: Saccade-Fixation Reiteration with Mamba for Referring Image Segmentation
Zhenjie Mao, Yuhuan Yang, Chaofan Ma +4
Referring Image Segmentation (RIS) aims to segment the target object in an image given a natural language expression. While recent methods leverage pre-trained vision backbones and…
SAM2MOT: A Novel Paradigm of Multi-Object Tracking by Segmentation
Junjie Jiang, Zelin Wang, Manqi Zhao +2
Inspired by Segment Anything 2, which generalizes segmentation from images to videos, we propose SAM2MOT--a novel segmentation-driven paradigm for multi-object tracking that breaks…
TerraGen: A Unified Multi-Task Layout Generation Framework for Remote Sensing Data Augmentation
Datao Tang, Hao Wang, Yudeng Xin +5
Remote sensing vision tasks require extensive labeled data across multiple, interconnected domains. However, current generative data augmentation frameworks are task-isolated, i.e.…
FIX-CLIP: Dual-Branch Hierarchical Contrastive Learning via Synthetic Captions for Better Understanding of Long Text
Bingchao Wang, Zhiwei Ning, Jianyu Ding +5
CLIP has shown promising performance across many short-text tasks in a zero-shot manner. However, limited by the input length of the text encoder, CLIP struggles on under-stream ta…