1 citations · 3 across the 5 of their papers we have counts for
6 papers
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
Weili Xu, Enxin Song, Wenhao Chai +3
The challenge of long video understanding lies in its high computational complexity and prohibitive memory cost, since the memory and computation required by transformer-based LLMs…
STEVE Series: Step-by-Step Construction of Agent Systems in Minecraft
Zhonghan Zhao, Wenhao Chai, Xuan Wang +7
Building an embodied agent system with a large language model (LLM) as its core is a promising direction. Due to the significant costs and uncontrollable factors associated with de…
MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
Enxin Song, Wenhao Chai, Tian Ye +3
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet…
VersaT2I: Improving Text-to-Image Models with Versatile Reward
Jianshu Guo, Wenhao Chai, Jie Deng +6
Recent text-to-image (T2I) models have benefited from large-scale and high-quality data, demonstrating impressive performance. However, these T2I models still struggle to produce i…
Hierarchical Auto-Organizing System for Open-Ended Multi-Agent Navigation
Zhonghan Zhao, Kewei Chen, Dongxu Guo +4
Due to the dynamic and unpredictable open-world setting, navigating complex environments in Minecraft poses significant challenges for multi-agent systems. Agents must interact wit…
See and Think: Embodied Agent in Virtual Environment
Zhonghan Zhao, Wenhao Chai, Xuan Wang +5
Large language models (LLMs) have achieved impressive pro-gress on several open-world tasks. Recently, using LLMs to build embodied agents has been a hotspot. This paper proposes S…