10 papers
Heterogeneity-Aware Belief Synchronization for Semantic Communication in AI-Native 6G Networks
Muhammad Hannan Akram, Muhammad Abubakar Rashid, Wassi Haider Kabir +3
6G networks will not be serving as communication infrastructures only; rather, they are expected to evolve into intelligent systems, where thousands of autonomous artificial intell…
When Do Fewer Visual Tokens Accelerate Multimodal Inference? A Break-Even Study Across Decision Locations and Hardware
Hao Dou, Ruiwen Tian
Fewer visual tokens do not guarantee lower end-to-end latency. We evaluate break-even with a reproducible protocol that accounts for decision overhead, shared work, and the operato…
WorldScape Policy 2.0: Empowering Steerable World Action Modeling with Reasoning-Augmented Memory
Haisheng Su, Zongdai Liu, Xin Jin +13
World Action Models (WAMs) offer a promising paradigm for robotic manipulation by jointly modeling visual state transitions and robot actions. However, existing WAMs are constraine…
CIGPO: Contextual Information-Gain Policy Optimization for Multi-Turn Evidence-Reading LLM Agents
Hao Dou
Training multi-turn evidence-reading agents with outcome-only reinforcement learning is unstable because intermediate turns receive little direct credit. In HotpotQA experiments wi…
DreamX-World 1.0: A General-Purpose Interactive World Model
DreamX Team, Yancheng Bai, Rui Chen +20
DreamX-World 1.0 is a general-purpose interactive text/image-to-video world model for controllable long-horizon generation. It supports camera navigation, revisits to previously ob…
RISE: Reliable Improvement in Self-Evolving Vision-Language Models
Chaoran Xu, Yingmao Miao, Pengfei Zhang +3
Vision-language models (VLMs) have achieved strong multimodal reasoning capabilities, but further improving them still relies heavily on large-scale human-constructed supervision f…