2 papers
cs.LG2026
From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation
Shuchang Ye, Jinqiang Yu, Zhujun Xiao +6
Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation t…
cs.CV2025
Exploring the Role of Explicit Temporal Modeling in Multimodal Large Language Models for Video Understanding
Yun Li, Zhe Liu, Yajing Kong +6
Applying Multimodal Large Language Models (MLLMs) to video understanding presents significant challenges due to the need to model temporal relations across frames. Existing approac…