22 citations · 55 across the 112 of their papers we have counts for
51 papers · 1 filter
Spatio-Temporal Fusion Model for Standard View Classification of Echocardiographic Videos
Bo Gou, Jicheng Zhang, Jianlong Xiong +9
Automated classification of standard echocardiographic views is crucial for efficient clinical workflow but faces three main challenges. First, publicly available datasets are scar…
AnchorWorld: Embodied Egocentric World Simulation with View-based Evolution Customization
Yu Li, Menghan Xia, Gongye Liu +8
Despite being a pivotal frontier, interactive world modeling remains underexplored in terms of the versatile controllability required by practical scenarios. To bridge this gap, we…
VEN-VL: A Visual Ensemble MoE Framework for Effective and Efficient Multi-Modal Understanding
Yinghao Wu, Zhuoyan Luo, Yiyao Yu +3
Despite the remarkable progress achieved by recent efficient methods in accelerating multimodal understanding, they still suffer from noticeable performance degradation. Their emph…
Video-Zero: Self-Evolution Video Understanding
Ruixu Zhang, Deyi Ji, Lanyun Zhu +4
Self-evolution offers a promising path for improving reasoning models without relying on intensive human annotation. However, extending this paradigm to video understanding remains…
IG-Diff: Complex Night Scene Restoration with Illumination-Guided Diffusion Model
Yifan Chen, Fei Yin, Chunle Guo +2
In nighttime circumstances, it is challenging for individuals and machines to perceive their surroundings. While prevailing image restoration methods adeptly handle singular forms…
MMCL-Bench: Multimodal Context Learning from Visual Rules, Procedures, and Evidence
Yifan Chen, Fei Yin, Qingyan Bai +2
We introduce MMCL-Bench, a benchmark for multimodal context learning: learning task-local rules, procedures, and empirical patterns from visual or mixed-modality teaching context a…