2 citations · 5 across the 3 of their papers we have counts for
3 papers
cs.CL2025★ 2 cited
Qwen3-Omni Technical Report
Jin Xu, Zhifang Guo, Hangrui Hu +35
We present Qwen3-Omni, a single multimodal model that, for the first time, maintains state-of-the-art performance across text, image, audio, and video without any degradation relat…
cs.MM2025★ 1 cited
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
Baoquan Zhao, Xiaofan Ma, Qianshi Pang +3
The widespread adoption of digital technology has ushered in a new era of digital transformation across all aspects of our lives. Online learning, social, and work activities, such…
cs.CV2025★ 2 cited
DepthGait: Multi-Scale Cross-Level Feature Fusion of RGB-Derived Depth and Silhouette Sequences for Robust Gait Recognition
Xinzhu Li, Juepeng Zheng, Yikun Chen +7
Robust gait recognition requires highly discriminative representations, which are closely tied to input modalities. While binary silhouettes and skeletons have dominated recent lit…