causal pretraining 1foundation models 1robot control 1sparse mixture of experts 1video-action models 1
From the 1 of 9 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
How Foundational Skills Influence VLM-based Embodied Agents:A Native Perspective
Bo Peng, Pi Bu, Keyu Pan +7
Recent advances in vision-language models (VLMs) have shown promise for human-level embodied intelligence. However, existing benchmarks for VLM-driven embodied agents often rely on…
cs.AI2025
YETI (YET to Intervene) Proactive Interventions by Multimodal AI Agents in Augmented Reality Tasks
Saptarashmi Bandyopadhyay, Vikas Bahirwani, Lavisha Aggarwal +3
Multimodal AI Agents are AI models that have the capability of interactively and cooperatively assisting human users to solve day-to-day tasks. Augmented Reality (AR) head worn dev…