3 papers
cs.CV2026
SAM3-I: Segment Anything with Instructions
Jingjing Li, Yue Feng, Yuchen Guo +10
Segment Anything Model 3 (SAM3) advances open-vocabulary segmentation through promptable concept segmentation, enabling users to segment all instances associated with a given conce…
cs.SD2025
Semantics-Aware Human Motion Generation from Audio Instructions
Zi-An Wang, Shihao Zou, Shiyao Yu +2
Recent advances in interactive technologies have highlighted the prominence of audio signals for semantic encoding. This paper explores a new task, where audio signals are used as…
cs.CV2025
SpikeVideoFormer: An Efficient Spike-Driven Video Transformer with Hamming Attention and Complexity
Shihao Zou, Qingfeng Li, Wei Ji +4
Spiking Neural Networks (SNNs) have shown competitive performance to Artificial Neural Networks (ANNs) in various vision tasks, while offering superior energy efficiency. However,…