From the 1 of 5 linked papers with an AI index.
5 papers
EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection
Wenhao Zhang, Kuanwei Lin, Xuyi Yang +2
The paper introduces EFlow, a framework that first retrieves visual evidence from long videos before reasoning, using separate chain‑of‑thought modules for temporal grounding and a…
EvalVerse: Pipeline-Aware and Expert-Calibrated Benchmarking for Professional Cinematic Video Generation
Songlin Yang, Haobin Zhong, Ruilin Zhang +23
The rapid evolution of generative video foundation models has propelled the field toward professional-grade cinematic synthesis. To achieve such demanding quality, the community tr…
ShotVerse: Advancing Cinematic Camera Control for Text-Driven Multi-Shot Video Creation
Songlin Yang, Zhe Wang, Xuyi Yang +7
Text-driven video generation has democratized film creation, but camera control in cinematic multi-shot scenarios remains a significant block. Implicit textual prompts lack precisi…
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
Xuyi Yang, Wenhao Zhang, Hongbo Jin +5
Current Multimodal Large Language Models (MLLMs) often perform poorly in long video understanding, primarily due to resource limitations that prevent them from processing all video…
CODESYNC: Synchronizing Large Language Models with Dynamic Code Evolution at Scale
Chenlong Wang, Zhaoyang Chu, Zhengxiang Cheng +6
Large Language Models (LLMs) have exhibited exceptional performance in software engineering yet face challenges in adapting to continually evolving code knowledge, particularly reg…