collaborators

7 papers

cs.CV2026

World Tokens: Enhancing Embodied Policies with Training-Time World Modeling

Qu Tang, Benhui Zhuang, Bo Yuan +3

Vision-language-action (VLA) models are a widely adopted paradigm for embodied policies. They excel at efficient closed-loop control but do not explicitly model how physical scenes…

cs.AI2026

SeerGuard: A Safety Framework for Mobile GUI Agents via World Model Prediction

Xue Yu, Bo Yuan, Pengshuai Yang +3

Mobile graphical user interface (GUI) agents have demonstrated remarkable capabilities in automating complex tasks, yet they introduce critical safety risks where a single erroneou…

cs.AI2026

JT-SAFE-V2: Safety-by-Design Foundation Model with World-Context Data

Junlan Feng, Fanyu Meng, Chong Long +12

We introduce JT-Safe-V2, a large language model designed to advance the safety and trustworthiness of foundation models, extending our previous JT-Safe model toward a more comprehe…

cs.CV2026

REC-RL: Referring expression counting via Gaussian and range-based reward optimization

Hui Liu, Yunlai Teng, Kunlong Bai +4

Referring expression counting (REC) is an intention-driven task that requires context-aware visual reasoning. While recent vision-language models incorporate language for visual un…

cs.CV2026

Tuning-free Instruction-based Video Editing Via Structural Noise Initialization and Guidance

Song Wu, Xinyu Chen, Qian Wang +3

Video editing poses a significant challenge. While a series of tuning-free methods circumvent the need for extensive data collection and model training, they often underutilize the…

cs.CV2026

GeoTikzBridge: Advancing Multimodal Code Generation for Geometric Perception and Reasoning

Jiayin Sun, Caixia Sun, Boyu Yang +7

Multimodal Large Language Models (MLLMs) have recently demonstrated remarkable perceptual and reasoning abilities. However, they struggle to perceive fine-grained geometric structu…