5 papers
Claw-Anything: Benchmarking Always-On Personal Assistants with Broader Access to User's Digital World
Yusong Lin, Xinyuan Liang, Haiyang Wang +8
Large language model agents are increasingly envisioned as always-on personal assistants with access to anything relevant in the user's digital world. Yet current systems operate o…
Depth Adaptive Efficient Visual Autoregressive Modeling
Chunliang Li, Tianze Cao, Sanyuan Zhao
Visual Autoregressive (VAR) modeling inefficiently applies a fixed computational depth to each position when generating high-resolution images. While existing methods accelerate in…
CLI-Gym: Scalable CLI Task Generation via Agentic Environment Inversion
Yusong Lin, Haiyang Wang, Shuzhe Wu +4
Agentic coding requires agents to effectively interact with runtime environments, e.g., command line interfaces (CLI), so as to complete tasks like resolving dependency issues, fix…
World knowledge-enhanced Reasoning Using Instruction-guided Interactor in Autonomous Driving
Mingliang Zhai, Cheng Li, Zengyuan Guo +7
The Multi-modal Large Language Models (MLLMs) with extensive world knowledge have revitalized autonomous driving, particularly in reasoning tasks within perceivable regions. Howeve…
RepVF: A Unified Vector Fields Representation for Multi-task 3D Perception
Chunliang Li, Wencheng Han, Junbo Yin +2
Concurrent processing of multiple autonomous driving 3D perception tasks within the same spatiotemporal scene poses a significant challenge, in particular due to the computational…