2 papers
cs.LG2026
From Failure Taxonomy to Intervention: A Diagnostic Methodology for Industry-Scale AVLM in Video and Live-Streaming Platform Moderation
Shuchang Ye, Jinqiang Yu, Zhujun Xiao +6
Industry-scale video and live-streaming moderation imposes requirements that are difficult to satisfy with generic pretrained public models or external APIs, including adaptation t…
cs.CV2026
Event-VLA: Action-Conditioned Event Fusion for Robust Vision-Language-Action Model
Jiaxin Liu, Xun Xu, Zhenhao Zhang +5
Vision-Language-Action (VLA) models have become an important paradigm of embodied AI. However, existing VLA models typically assume well-lit and stable indoor settings, while real-…