3 papers
cs.AI2026
WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
Yehang Zhang, Jianchong Su, Haojian Huang +7
To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Existing long-term memory benchmarks mainly…
cs.CV2026
VideoMemory: Toward Consistent Video Generation via Memory Integration
Jinsong Zhou, Yihua Du, Xinli Xu +7
Maintaining consistent characters, props, and environments across multiple shots is a central challenge in narrative video generation. Existing models can produce high-quality shor…
cs.CV2025
Long-Video Audio Synthesis with Multi-Agent Collaboration
Yehang Zhang, Xinli Xu, Xiaojie Xu +2
Video-to-audio synthesis, which generates synchronized audio for visual content, critically enhances viewer immersion and narrative coherence in film and interactive media. However…