3 citations · 3 across the 5 of their papers we have counts for
5 papers
Spin decoherence dynamics of Er in CeO film
Sagar Kumar Seth, Jonah Nagura, Vrindaa Somjit +10
Developing telecom-compatible spin-photon interfaces is essential towards scalable quantum networks. Erbium ions (Er) exhibit a unique combination of a telecom (1.5 m) op…
VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking
Desen Meng, Rui Huang, Zhilin Dai +8
While recent advances in reinforcement learning have significantly enhanced reasoning capabilities in large language models (LLMs), these techniques remain underexplored in multi-m…
Towards Open-Vocabulary Video Semantic Segmentation
Xinhao Li, Yun Liu, Guolei Sun +3
Semantic segmentation in videos has been a focal point of recent research. However, existing models encounter challenges when faced with unfamiliar categories. To address this, we…
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
Xinhao Li, Zhenpeng Huang, Jing Wang +2
With the growth of high-quality data and advancement in visual pre-training paradigms, Video Foundation Models (VFMs) have made significant progress recently, demonstrating their r…
VideoMamba: State Space Model for Efficient Video Understanding
Kunchang Li, Xinhao Li, Yi Wang +4
Addressing the dual challenges of local redundancy and global dependencies in video understanding, this work innovatively adapts the Mamba to the video domain. The proposed VideoMa…