activity
20232025
most citedMovieChat+: Question-aware Sparse Memory for Long Video Question Answering

1 citations · 3 across the 5 of their papers we have counts for

collaborators

6 papers

cs.CV2025

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

Weili Xu, Enxin Song, Wenhao Chai +3

The challenge of long video understanding lies in its high computational complexity and prohibitive memory cost, since the memory and computation required by transformer-based LLMs…

cs.CV2024

STEVE Series: Step-by-Step Construction of Agent Systems in Minecraft

Zhonghan Zhao, Wenhao Chai, Xuan Wang +7

Building an embodied agent system with a large language model (LLM) as its core is a promising direction. Due to the significant costs and uncontrollable factors associated with de…

cs.CV20241 cited

MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Enxin Song, Wenhao Chai, Tian Ye +3

Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet…

cs.CV20241 cited

VersaT2I: Improving Text-to-Image Models with Versatile Reward

Jianshu Guo, Wenhao Chai, Jie Deng +6

Recent text-to-image (T2I) models have benefited from large-scale and high-quality data, demonstrating impressive performance. However, these T2I models still struggle to produce i…

cs.CV20241 cited

Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation

Zhonghan Zhao, Kewei Chen, Dongxu Guo +4

Due to the dynamic and unpredictable open-world setting, navigating complex environments in Minecraft poses significant challenges for multi-agent systems. Agents must interact wit…

cs.AI2023

See and Think: Embodied Agent in Virtual Environment

Zhonghan Zhao, Wenhao Chai, Xuan Wang +5

Large language models (LLMs) have achieved impressive pro-gress on several open-world tasks. Recently, using LLMs to build embodied agents has been a hotspot. This paper proposes S…