3 papers
cs.CV2026
MSEditor: Toward Consistent Multi-Shot Video Editing
Kunyu Feng, Yue Ma, Bingyuan Wang +6
In this paper, we tackle the problem of performing consistent, unified modifications to a multi-shot video sequence. This task is particularly challenging because multi-shot videos…
cs.RO2026
ST-BiBench: Benchmarking Multi-Stream Multimodal Coordination in Bimanual Embodied Tasks for MLLMs
Xin Wu, Zhixuan Liang, Yue Ma +3
Multimodal Large Language Models (MLLMs) have significantly advanced the landscape of embodied AI, yet transitioning to synchronized bimanual coordination introduces formidable cha…
cs.CV2025
Follow-Your-Instruction: A Comprehensive MLLM Agent for World Data Synthesis
Kunyu Feng, Yue Ma, Xinhua Zhang +9
With the growing demands of AI-generated content (AIGC), the need for high-quality, diverse, and scalable data has become increasingly crucial. However, collecting large-scale real…