1 citations · 2 across the 17 of their papers we have counts for
7 papers · 1 filter
Qwen-Image-2.0 Technical Report
Bing Zhao, Chenfei Wu, Deqing Li +72
We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite rece…
SpatialV2A: Visual-Guided High-fidelity Spatial Audio Generation
Yanan Wang, Linjie Ren, Zihao Li +2
While video-to-audio generation has achieved remarkable progress in semantic and temporal alignment, most existing studies focus solely on these aspects, paying limited attention t…
PA-VAD: Diffusion-Based Pseudo-Only Video Anomaly Detection via Domain-Aligned Memory Updates
Satoshi Hashimoto, Yanan Wang, Hitoshi Nishimura +1
Deploying video anomaly detection (VAD) in the real world is often constrained by the scarcity, privacy, and cost of collecting real abnormal footage. We propose PA-VAD, a novel ps…
MARS2 2025 Challenge on Multimodal Reasoning: Datasets, Methods, Results, Discussion, and Outlook
Peng Xu, Shengwu Xiong, Jiajun Zhang +125
This paper reviews the MARS2 2025 Challenge on Multimodal Reasoning. We aim to bring together different approaches in multimodal machine learning and LLMs via a large benchmark. We…
CoTasks: Chain-of-Thought based Video Instruction Tuning Tasks
Yanan Wang, Julio Vizcarra, Zhi Li +2
Despite recent progress in video large language models (VideoLLMs), a key open challenge remains: how to equip models with chain-of-thought (CoT) reasoning abilities grounded in fi…
Top-down Activity Representation Learning for Video Question Answering
Yanan Wang, Shuichiro Haruta, Donghuo Zeng +2
Capturing complex hierarchical human activities, from atomic actions (e.g., picking up one present, moving to the sofa, unwrapping the present) to contextual events (e.g., celebrat…