7 papers
World Tokens: Enhancing Embodied Policies with Training-Time World Modeling
Qu Tang, Benhui Zhuang, Bo Yuan +3
Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes…
SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction
Xue Yu, Bo Yuan, Pengshuai Yang +3
Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneou…
JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data
Junlan Feng, Fanyu Meng, Chong Long +12
We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe model toward a more comprehe…
REC-RL: Referring expression counting via Gaussian and range-based reward optimization
Hui Liu, Yunlai Teng, Kunlong Bai +4
Referring expression counting (REC) is an intention-driven task that requires context-aware visual reasoning. While recent vision-language models incorporate language for visual un…
Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance
Song Wu, Xinyu Chen, Qian Wang +3
Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the…
GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning
Jiayin Sun, Caixia Sun, Boyu Yang +7
Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable perceptual and reasoning abilities. However, they struggle to perceive fine-grained geometric structu…