From the 1 of 15 linked papers with an AI index.
15 papers
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
Dwip Dalal, Shivansh Patel, Chahit Jain +7
The paper introduces Anchor-Align, a method that adds representation anchoring and language-action alignment to behavior‑cloning finetuning of vision‑language models for robot mani…
NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding
Hyeonjeong Ha, Jinjin Ge, Bo Feng +2
Multimodal large language models (MLLMs) have achieved impressive progress in vision-language reasoning, yet their ability to understand temporally unfolding narratives in videos r…
Trimming the Long-Tail of Visual World Modeling Evaluation
Bingxuan Li, Yining Hong, Cheng Qian +6
Physical interactions follow a long-tailed distribution: a set of common and regular interactions dominates human experience and visual data, while a broad spectrum of rare and irr…
Advancing Creative Physical Intelligence in Large Multimodal Models
Cheng Qian, Hyeonjeong Ha, Jiayu Liu +10
Large multimodal models (LMMs) have rapidly advanced in perception and reasoning; however, it remains unclear whether these capabilities generalize to discovering visually grounded…
MemGuard: Preventing Memory Contamination in Long-Term Memory-Augmented Large Language Models
Hyeonjeong Ha, Jeonghwan Kim, Cheng Qian +7
Memory-augmented large language models extend reasoning beyond a fixed context window by maintaining long-term memory across interactions. However, existing memory systems often co…
MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Poisoning Attacks
Hyeonjeong Ha, Qiusi Zhan, Jeonghwan Kim +6
Retrieval-augmented generation (RAG) has become a common practice in multimodal large language models (MLLM) to enhance factual grounding and reduce hallucination. Yet, its relianc…