3 papers
cs.CV2026
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
Chaojie Mao, Chen-Wei Xie, Chongyang Zhong +55
We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivi…
cs.CV2026
Reinforcing Structured Chain-of-Thought for Video Understanding
Peiyao Wang, Haotian Xu, Noranart Vesdapunt +6
Multi-modal Large Language Models (MLLMs) show promise in video understanding. However, their reasoning often suffers from thinking drift and weak temporal comprehension, even when…
cs.IR2025
FG-RAG: Enhancing Query-Focused Summarization with Context-Aware Fine-Grained Graph RAG
Yubin Hong, Chaofan Li, Jingyi Zhang +1
Retrieval-Augmented Generation (RAG) enables large language models to provide more precise and pertinent responses by incorporating external knowledge. In the Query-Focused Summari…