1 citations · 1 across the 12 of their papers we have counts for
21 papers
MemWM: Memory-Augmented Text-Based World Model
Yujun Wang, Tao Zhang, Jinhe Bi +9
World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can sti…
DriveCache: Action-Aware Caching for Driving World Model Inference
Jianchun Yang, Jian Liang, Xianda Guo +5
Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Di…
Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models
Wenwen He, Wenke Huang, Wei Yang Bryan Lim +1
Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and humans. Yet it remains unclear whether their influence is driven by persuasion-o…
ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning
Jinhe Bi, Chennan Zhou, Zengjie Jin +10
On-policy training has emerged as a powerful post-training paradigm for improving the reasoning capabilities of large language models, and is often enhanced by golden trajectories…
Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning
Qianlong Yang, Bowen Ye, Xianda Guo +4
Despite the progress of multimodal large language models (MLLMs), they continue to exhibit deficiencies in visual perception. Following visual instruction tuning, internal MLLM rep…
TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors
Pinhan Fu, Xianda Guo, Xuetao Li +5
Vision-Language-Action (VLA) models are deployed through pipelines that end users cannot audit, and a poisoned VLA can behave normally on clean observations while a small visual tr…