2 papers
cs.CV2025
StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification
Yichen He, Yuan Lin, Jianchao Wu +3
Existing large vision-language models (LVLMs) are largely limited to processing short, seconds-long videos and struggle with generating coherent descriptions for extended video spa…
cs.LG2024
AGILE: A Novel Reinforcement Learning Framework of LLM Agents
Peiyuan Feng, Yichen He, Guanhua Huang +4
We introduce a novel reinforcement learning framework of LLM agents named AGILE (AGent that Interacts and Learns from Environments) designed to perform complex conversational tasks…