3 papers
cs.RO2026
AnyScene: Towards Highly Controllable Driving Scene Generation at Anywhere and Beyond
Haiming Zhang, Junfei Zhou, Feng Jiang +6
Generating high-fidelity and controllable synthetic data is critical for advancing end-to-end autonomous driving, particularly for addressing the long tail of rare safety-critical…
cs.CV2026
Divide-and-Conquer Inference for Large-Scale Visual Recognition with Multimodal Large Language Models
Zhipeng Ye, Jiaqi Huang, Feng Jiang +5
Multimodal Large Language Models (MLLMs) have demonstrated strong capabilities across a wide range of vision language tasks. However, when applied to large scale image classificati…
cs.CV2026
AoE: Always-on Egocentric Human Video Collection for Embodied AI
Bowen Yang, Zishuo Li, Yang Sun +15
Embodied foundation models require large-scale, high-quality real-world interaction data for pre-training and scaling. However, existing data collection methods suffer from high in…