Publications (8)
StoryTeller: Improving Long Video Description through Global Audio-Visual Character Identification
Yichen He, Yuan Lin, Jianchao Wu +3
Existing large vision-language models (LVLMs) are largely limited to processing short, seconds-long videos and struggle with generating coherent descriptions for extended video spa…
Delving into Commit-Issue Correlation to Enhance Commit Message Generation Models
Liran Wang, Xunzhu Tang, Yichen He +4
Commit message generation (CMG) is a challenging task in automated software engineering that aims to generate natural language descriptions of code changes for commits. Previous me…
PaSa: An LLM Agent for Comprehensive Academic Paper Search
Yichen He, Guanhua Huang, Peiyuan Feng +4
We introduce PaSa, an advanced Paper Search agent powered by large language models. PaSa can autonomously make a series of decisions, including invoking search tools, reading paper…
Seeing, Listening, Remembering, and Reasoning: A Multimodal Agent with Long-Term Memory
Lin Long, Yichen He, Wentao Ye +5
We introduce M3-Agent, a novel multimodal agent framework equipped with long-term memory. Like humans, M3-Agent can process real-time visual and auditory inputs to build and update…
AGILE: A Novel Reinforcement Learning Framework of LLM Agents
Peiyuan Feng, Yichen He, Guanhua Huang +4
We introduce a novel reinforcement learning framework of LLM agents named AGILE (AGent that Interacts and Learns from Environments) designed to perform complex conversational tasks…
FraudBench: A Multimodal Benchmark for Detecting AI-Generated Fraudulent Refund Evidence
Xinyu Yan, Boyang Chen, Jiaming Zhang +12
Artificial Intelligence (AI)-generated images have become increasingly realistic and readily adaptable to concrete real-world claims, creating new challenges for verifying visual e…
Feature Selection in High-dimensional Spaces Using Graph-Based Methods
Swarnadip Ghosh, Somabha Mukherjee, Divyansh Agarwal +3
High-dimensional feature selection is a central problem in a variety of application domains such as machine learning, image analysis, and genomics. In this paper, we propose graph-…
Task-Focused Memorization for Multimodal Agents
Tao Zou, Yichen He, Tian Qiu +2
Long-term memory is essential for multimodal agents to build coherent experience, accumulate world knowledge, and achieve continual learning. However, constructing effective memory…