4 citations · 8 across the 17 of their papers we have counts for
Showing 2024 · cs.CVShow all
3 papers · 2 filters
cs.CV2024
RoomTour3D: Geometry-Aware Video-Instruction Tuning for Embodied Navigation
Mingfei Han, Liang Ma, Kamila Zhumakhanova +5
Vision-and-Language Navigation (VLN) suffers from the limited diversity and scale of training data, primarily constrained by the manual curation of existing simulators. To address…
cs.CV2024
StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration
Panwen Hu, Jin Jiang, Jianqi Chen +4
The advent of AI-Generated Content (AIGC) has spurred research into automated video generation to streamline conventional processes. However, automating storytelling video producti…
cs.CV2024★ 1 cited
LongVLM: Efficient Long Video Understanding via Large Language Models
Yuetian Weng, Mingfei Han, Haoyu He +2
Empowered by Large Language Models (LLMs), recent advancements in Video-based LLMs (VideoLLMs) have driven progress in various video understanding tasks. These models encode video…