From the 1 of 24 linked papers with an AI index.
24 papers
DriveCache: Action-Aware Caching for Driving World Model Inference
Jianchun Yang, Jian Liang, Xianda Guo +5
Driving video generation models support autonomous-driving development by predicting controllable future scenes for simulation, planning evaluation, and offline data generation. Di…
MemWM: Memory-Augmented Text-Based World Model
Yujun Wang, Tao Zhang, Jinhe Bi +9
World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can sti…
TrustVLA: Mechanism-Guided Inference-Time Defense Against Vision-Language-Action Backdoors
Pinhan Fu, Xianda Guo, Xuetao Li +5
The paper introduces TrustVLA, an inference-time defense that detects and mitigates visual backdoor triggers in vision‑language‑action models by monitoring epistemic uncertainty an…
Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models
Xuankun Rong, Wenke Huang, Bo Du +2
As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable be…
Switch-Reasoner: Learn When to Think in Multitask Mixtures via Reinforcement Learning
Yiyang Fang, Pei Fu, Jinjie Li +7
Multimodal Large Language Models (MLLMs) often follow a fixed Think-then-Answer paradigm, which is inefficient in heterogeneous multitask settings because simple inputs may not req…
EMO-R3: Reflective Reinforcement Learning for Emotional Reasoning in Multimodal Large Language Models
Yiyang Fang, Wenke Huang, Pei Fu +5
Multimodal Large Language Models (MLLMs) have shown remarkable progress in visual reasoning and understanding tasks but still struggle to capture the complexity and subjectivity of…