activity
20222026
most citedEfficient Multimodal Large Language Models: A Survey

32 citations · 94 across the 24 of their papers we have counts for

collaborators
Showing cs.CVShow all

25 papers · 1 filter

cs.CV2026

VideoSearcher: Empowering Video Deep Research with Multi-Tool Agentic Reasoning via Reinforcement Learning

Zhenkun Gao, Yicheng Bao, Jinlong Peng +13

Video understanding is moving beyond closed-context perception toward open-world evidence exploration, a paradigm formalized as Video Deep Research (VDR). However, existing multimo…

cs.CV20253 cited

YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection

Xu Lin, Jinlong Peng, Zhenye Gan +2

Existing Real-Time Object Detection (RTOD) methods commonly adopt YOLO-like architectures for their favorable trade-off between accuracy and speed. However, these models rely on st…

cs.CV2025

Soul: Breathe Life into Digital Human for High-fidelity Long-term Multimodal Animation

Jiangning Zhang, Junwei Zhu, Zhenye Gan +14

We propose a multimodal-driven framework for high-fidelity long-term digital human animation termed , which generates semantically coherent videos from a single-fram…

cs.CV2025

Transform Trained Transformer: Accelerating Naive 4K Video Generation Over 10

Jiangning Zhang, Junwei Zhu, Teng Hu +7

Native 4K (21603840) video generation remains a critical challenge due to the quadratic computational explosion of full-attention as spatiotemporal resolution increases, ma…

cs.CV2025

Real-IAD Variety: Pushing Industrial Anomaly Detection Dataset to a Modern Era

Wenbing Zhu, Chengjie Wang, Bin-Bin Gao +12

Industrial Anomaly Detection (IAD) is a cornerstone for ensuring operational safety, maintaining product quality, and optimizing manufacturing efficiency. However, the advancement…

cs.CV2025

Human-MME: A Holistic Evaluation Benchmark for Human-Centric Multimodal Large Language Models

Yuansen Liu, Haiming Tang, Jinlong Peng +12

Multimodal Large Language Models (MLLMs) have demonstrated significant advances in visual understanding tasks. However, their capacity to comprehend human-centric scenes has rarely…