From the 1 of 46 linked papers with an AI index.
46 papers
ForeDreamer: A Self-Evolving Dual-Agent Memory Architecture for Future Event Prediction
Linhao Zhong, Zongze Du, Linyu Wu +6
Open-web future event prediction requires agents to distill reliable signals from noisy, redundant, and incomplete evidence. Existing retrieval/memory mechanisms directly feed retr…
Open-AoE: An Open Egocentric Manipulation Dataset and Toolchain for Embodied Learning
Zishuo Li, Bowen Yang, Changtao Miao +29
The paper introduces Open-AoE, a large-scale egocentric video dataset of human manipulation together with a processing and downstream toolchain that provides hand pose, camera traj…
CoRE-VLA: Towards Scalable and Robust Vision-Language-Action Modeling via Conditional Routing of Experts
Haozhe Zhang, Sixian Li, Yifei Zhang +5
Vision-language-action (VLA) models have advanced generalist robotic manipulation, yet real-world deployment reveals a fundamental challenge: robots are equipped with diverse and h…
MemDreamer: Decoupling Perception and Reasoning for Long Video Understanding via Hierarchical Graph Memory and Agentic Retrieval Mechanism
Cong Chen, Guo Gan, Kaixiang Ji +7
Current Vision-Language Models struggle with hours-long videos because processing full-length visual sequences induces prohibitive token explosion and attention dilution. To overco…
GAE: Unleashing Physical Potential of VLM with Generalizable Action Expert
Mingyu Liu, Zheng Huang, Xiaoyi Lin +6
Vision-language models demonstrate strong reasoning and planning abilities, yet grounding these predictions into precise robot actions remains a central challenge. Existing Vision-…
ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning
Muzhi Zhu, Hao Zhong, Canyu Zhao +9
Active vision, also known as active perception, refers to actively selecting where and how to look in order to gather task-relevant information. It is a critical component of effic…