From the 1 of 14 linked papers with an AI index.
14 papers
MetaVideoAgent: Automated Video-Agent Evolution for Long-Form Video Understanding
Benlei Cui, Ruize Wang, Junjie Li +9
Long-form video understanding requires locating sparse, question-relevant evidence in long, multimodal videos. Real-world video distributions differ in modality-specific informatio…
Pass the Baton: Trajectory-Relayed On-Policy Distillation
Haolei Xu, Xiaowen Xu, Haiwen Hong +5
The paper proposes Relay On-Policy Distillation (Relay-OPD), a method that lets a teacher model temporarily take over a student’s generation when a wrong reasoning prefix is detect…
YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models
Junyu Lin, Meizhen Liu, Xiufeng Huang +12
As large language models (LLMs) are increasingly deployed in real-world applications, safety guardrails are required to go beyond coarse-grained filtering and support fine-grained,…
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning
Hongxing Li, Xiufeng Huang, Dingming Li +11
Fine-grained visual reasoning remains challenging for vision-language models, especially when small but critical visual cues are buried in high-resolution images. Existing approach…
Yuvion LLM: An Adversarially-Aware Large Language Model for Content And AI Safety
Ting Ma, Xiufeng Huang, Benlei Cui +43
As large language models are increasingly deployed in real-world systems, safety failures can still lead to harmful outputs and dangerous misuse. We argue that the essence of safet…
Yuvion VL: A Multimodal Foundation Model for Adversarial Content and AI Safety
Shikai Qiu, Xiaowen Xu, Benlei Cui +55
General-purpose models often struggle to reliably identify and understand real-world multimodal risks, largely due to the inherent multimodal adversarial nature of content and AI s…