4 papers
MemWM: Memory-Augmented Text-Based World Model
Yujun Wang, Tao Zhang, Jinhe Bi +9
World models are increasingly used to support planning in agents by predicting how environment states evolve in response to agent actions. Yet fluent next-state predictions can sti…
FabriMAE I Trust Myself? Self-Evaluating VLA Action Generation with Markov Attention Entropy
Aniri, Chen Yilin, Jinhe Bi +10
Vision-Language-Action models (VLAs) integrate visual perception, language instruction, and action generation into end-to-end policies across heterogeneous architectures. However,…
CausalMix: Data Mixture as Causal Inference for Language Model Training
Zinan Tang, Yukun Zhang, Shaomian Zheng +6
In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely o…
LLaVA Steering: Visual Instruction Tuning with 500x Fewer Parameters through Modality Linear Representation-Steering
Jinhe Bi, Yujun Wang, Haokun Chen +4
Multimodal Large Language Models (MLLMs) have significantly advanced visual tasks by integrating visual representations into large language models (LLMs). The textual modality, inh…