5 papers
BridgeV2W: Bridging Video Generation Models to Embodied World Models via Embodiment Masks
Yixiang Chen, Peiyan Li, Jiabing Yang +8
Embodied world models have emerged as a promising paradigm in robotics, most of which leverage large-scale Internet videos or pretrained video generation models to enrich visual an…
ToolWeaver: Weaving Collaborative Semantics for Scalable Tool Use in Large Language Models
Bowen Fang, Wen Ye, Yunyue Su +8
Prevalent retrieval-based tool-use pipelines struggle with a dual semantic challenge: their retrievers often employ encoders that fail to capture complex semantics, while the Large…
BrowserAgent: Building Web Agents with Human-Inspired Web Browsing Actions
Tao Yu, Zhengbo Zhang, Zhiheng Lyu +8
Efficiently solving real-world problems with LLMs increasingly hinges on their ability to interact with dynamic web environments and autonomously acquire external information. Whil…
IKOD: Mitigating Visual Attention Degradation in Large Vision-Language Models
Jiabing Yang, Chenhang Cui, Yiyang Zhou +6
Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated significant progress across multiple domains. However, these models still face the inherent challenge…
EC-Flow: Enabling Versatile Robotic Manipulation from Action-Unlabeled Videos via Embodiment-Centric Flow
Yixiang Chen, Peiyan Li, Yan Huang +3
Current language-guided robotic manipulation systems often require low-level action-labeled datasets for imitation learning. While object-centric flow prediction methods mitigate t…