4 papers
Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning
Shiding Zhu, Yudi Qi, Yajie Wang +6
Experience-driven self-evolution is critical for large language model (LLM) agents to improve through open-world interaction. However, existing experience learning methods mostly r…
FILA: Fine-Grained Vision Language Models
Shiding Zhu, Wenhui Dong, Jun Song +3
Recently, there has been growing interest in the capability of multimodal large language models (MLLMs) to process high-resolution images. A common approach currently involves dyna…
FiLA-Video: Spatio-Temporal Compression for Fine-Grained Long Video Understanding
Yanan Guo, Wenhui Dong, Jun Song +7
Recent advancements in video understanding within visual large language models (VLLMs) have led to notable progress. However, the complexity of video data and contextual processing…
Multilevel Semantic-Aware Model for AI-Generated Video Quality Assessment
Jiaze Li, Haoran Xu, Shiding Zhu +2
The rapid development of diffusion models has greatly advanced AI-generated videos in terms of length and consistency recently, yet assessing AI-generated videos still remains chal…