7 papers
In-Context VLA: Endowing Vision-Language-Action Models with Language via In-Context Post-Training and Agentic Tool Use
Jiarui Yang, Wen Huang, Jiale Zhang +2
Vision-Language-Action (VLA) models have become the dominant recipe for generalist manipulation, yet they are almost universally trained by behavior cloning: a policy imitates expe…
LAWM-3D: Learning 3D-Aware Latent Actions from Human Videos for Generalizable Robot World Models
Jiarui Yang, Jiale Zhange, Jiawei Li +5
World models enable agents to perform forward rollout and planning without real-world interaction. However, their application in open-world embodied intelligence remains limited by…
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
Wen Huang, Jiarui Yang, Tao Dai +4
Visual manipulation localization (VML) aims to identify tampered regions in images and videos, a task that has become increasingly challenging with the rise of advanced editing too…
ARBOR: Online Process Rewards via a Reusable Rubric Buffer for Search Agents
Zheng Liu, Longxiang Zhang, Xintong Wang +8
LLM-based search agents are trained predominantly with outcome-only reward, leaving the search process itself unsupervised. This signal degenerates on outcome-homogeneous groups wh…
BraveGuard: From Open-World Threats to Safer Computer-Use Agents
Yunhao Feng, Xiaohu Du, Xinhao Deng +13
Computer-use agents extend language models from text generation to sustained interaction with files, terminals, browsers, and external tools. This shift creates safety risks that a…
Improving Deepfake Detection with Reinforcement Learning-Based Adaptive Data Augmentation
Yuxuan Zhou, Tao Yu, Wen Huang +3
The generalization capability of deepfake detectors is critical for real-world use. Data augmentation via synthetic fake face generation effectively enhances generalization, yet cu…