3 papers
cs.CV2026
ViPSim: Collaborating Visual and Parameter Spaces for Consistent Long-Horizon Embodied World Models
Longyu Chen, Heng Li, Wei Yang +2
Embodied World Models (EWMs) have emerged as a scalable and risk-free paradigm for advancing embodied intelligence, enabling the safety-critical evaluation of Vision-Language-Actio…
cs.CV2026
ReGLA: Efficient Receptive-Field Modeling with Gated Linear Attention Network
Junzhou Li, Manqi Zhao, Yilin Gao +4
Balancing accuracy and latency on high-resolution images is a critical challenge for lightweight models, particularly for Transformer-based architectures that often suffer from exc…
cs.CV2025
SAM2MOT: A Novel Paradigm of Multi-Object Tracking by Segmentation
Junjie Jiang, Zelin Wang, Manqi Zhao +2
Inspired by Segment Anything 2, which generalizes segmentation from images to videos, we propose SAM2MOT--a novel segmentation-driven paradigm for multi-object tracking that breaks…