works on

From the 1 of 11 linked papers with an AI index.

activity
20242026
collaborators

11 papers

cs.AI2026

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

Chunyang Jiang, Pingping Zhang, Yuzhi Zhao +9

Multimodal large language models (MLLMs) have achieved remarkable performance across vision-language tasks, but their progress depends heavily on large-scale, high-quality multimod…

cs.RO2026

Zero2Skill: Bootstrapping Robot Skills through Autonomous Data Collection, Training, and Deployment

Boyuan Wang, Zhenyuan Zhang, Zhiqin Yang +16

Zero2Skill is a robot learning system that autonomously collects, verifies, and resets manipulation data while using a large language model to store and reuse human corrections, dr…

cs.LG2026

A Control Theory of Predictability in Latent World Models

Hanzhe You, Yonggang Zhang, Maohao Ran +6

Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward. Current…

cs.AI2026

MOSS: Self-Evolution through Source-Level Rewriting in Autonomous Agent Systems

Qianshu Cai, Yonggang Zhang, Xianzhang Jia +5

Autonomous agentic systems are largely static after deployment: they do not learn from user interactions, and recurring failures persist until the next human-driven update ships a…

cs.AI2026

Conditional Equivalence of DPO and RLHF: Implicit Assumption, Failure Modes, and Provable Alignment

Zhiqin Yang, Yonggang Zhang, Wei Xue +3

Direct Preference Optimization (DPO) has emerged as a popular alternative to Reinforcement Learning from Human Feedback (RLHF), offering theoretical equivalence with simpler implem…

cs.CV2026

Attention Hijacking: Response Manipulation Across Queries in Vision-Language Models

Zhiqiang Wang, Dongrui Liu, Yan Li +4

Existing adversarial attacks on vision-language models (VLMs) can steer model outputs toward attacker-specified target responses, but their effectiveness often degrades when the sa…