32 citations · 94 across the 24 of their papers we have counts for
25 papers · 1 filter
VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning
Zhenkun Gao, Yicheng Bao, Jinlong Peng +13
Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (VDR). However, existing multimo…
YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection
Xu Lin, Jinlong Peng, Zhenye Gan +2
Existing Real-Time Object Detection (RTOD) methods commonly adopt YOLO-like architectures for their favorable trade-off between accuracy and speed. However, these models rely on st…
Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation
Jiangning Zhang, Junwei Zhu, Zhenye Gan +14
We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed , which generates semantically coherent videos from a single-fram…
Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10
Jiangning Zhang, Junwei Zhu, Teng Hu +7
Native 4K (21603840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, ma…
Real-IAD Variety: Pushing Industrial Anomaly Detection Dataset to a Modern Era
Wenbing Zhu, Chengjie Wang, Bin-Bin Gao +12
Industrial Anomaly Detection (IAD) is a cornerstone for ensuring operational safety, maintaining product quality, and optimizing manufacturing efficiency. However, the advancement…
Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models
Yuansen Liu, Haiming Tang, Jinlong Peng +12
Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely…