most citedBeyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation

1 citations · 1 across the 4 of their papers we have counts for

collaborators

7 papers

cs.CV2026

Facial-R1: Aligning Reasoning and Recognition for Facial Emotion Analysis

Jiulong Wu, Yucheng Shen, Lingyong Yan +4

Facial Emotion Analysis (FEA) extends traditional facial emotion recognition by incorporating explainable, fine-grained reasoning. The task integrates three subtasks: emotion recog…

cs.AI20261 cited

Beyond End-to-End Video Models: An LLM-Based Multi-Agent System for Educational Video Generation

Lingyong Yan, Jiulong Wu, Dong Xie +3

Although recent end-to-end video generation models demonstrate impressive performance in visually oriented content creation, they remain limited in scenarios that require strict lo…

cs.CV2026

VISOR: Agentic Visual Retrieval-Augmented Generation via Iterative Search and Over-horizon Reasoning

Yucheng Shen, Jiulong Wu, Jizhou Huang +3

Visual Retrieval-Augmented Generation (VRAG) empowers Vision-Language Models to retrieve and reason over visually rich documents. To tackle complex queries requiring multi-step rea…

cs.AI2026

Adversarial Yet Cooperative: Multi-Perspective Reasoning in Retrieved-Augmented Language Models

Can Xu, Lingyong Yan, Jiayi Wu +6

Recent advances in synergizing large reasoning models (LRMs) with retrieval-augmented generation (RAG) have shown promising results, yet two critical challenges remain: (1) reasoni…

cs.CL2025

GRAF: Multi-turn Jailbreaking via Global Refinement and Active Fabrication

Hua Tang, Lingyong Yan, Yukun Zhao +3

Large Language Models (LLMs) have demonstrated remarkable performance across diverse tasks. Nevertheless, they still pose notable safety risks due to potential misuse for malicious…

cs.CV2025

Mitigating Hallucinations in Large Vision-Language Models via Entity-Centric Multimodal Preference Optimization

Jiulong Wu, Zhengliang Shi, Shuaiqiang Wang +5

Large Visual Language Models (LVLMs) have demonstrated impressive capabilities across multiple tasks. However, their trustworthiness is often challenged by hallucinations, which ca…