10 papers
Skill-Use: Can LLMs Actually Use Skills in Agentic Harnesses?
Jinyi Han, Yuanjian Xu, Ying Liao +6
Large language model (LLM) agents increasingly rely on skills, structured documents that specify when to act, which procedure to follow, and which tools are allowed. Existing evalu…
From Foundation to Application: Improving VLA Models in Practice
Wei Wu, Fangjing Wang, Fan Lu +21
Despite recent progress of VLA foundation models, the disparity between laboratory conditions and real-world applications continues to impede their practical implementation. To bri…
A Pragmatic VLA Foundation Model
Wei Wu, Fan Lu, Yunnan Wang +22
Offering great potential in robotic manipulation, a capable Vision-Language-Action (VLA) foundation model is expected to faithfully generalize across tasks and platforms while ensu…
VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes
Jingru Chen, Yiming Liu, Mingtao Chen +5
Frontier multimodal large language models (MLLMs) have been reported to achieve over 90% accuracy on fine-grained perception benchmarks. However, such scores do not necessarily imp…
Self-Consistent Latent Reasoning: Long Latent Sequence Reasoning for Vision-Language Model
Chenfeng Wang, Wei He, Xuhan Zhu +10
In language reasoning, longer chains of thought consistently yield better performance, which naturally suggests that visual latent reasoning may likewise benefit from longer latent…
Visually-grounded Humanoid Agents
Hang Ye, Xiaoxuan Ma, Fan Lu +3
Digital human generation has been studied for decades and supports a wide range of real-world applications. However, most existing systems are passively animated, relying on privil…