3 papers
cs.CV2026
Frozen LLMs as Map-Aware Spatio-Temporal Reasoners for Vehicle Trajectory Prediction
Yanjiao Liu, Jiawei Liu, Xun Gong +1
Large language models (LLMs) have recently demonstrated strong reasoning capabilities and attracted increasing research attention in the field of autonomous driving (AD). However,…
cs.MM2025
An Evaluation of Interleaved Instruction Tuning on Semantic Reasoning Performance in an Audio MLLM
Jiawei Liu, Enis Berk Ãoban, Zarina Schevchenko +4
Standard training for Multi-modal Large Language Models (MLLMs) involves concatenating non-textual information, like vision or audio, with a text prompt. This approach may not enco…
cs.CV2025
Beyond Semantics: Rediscovering Spatial Awareness in Vision-Language Models
Jianing Qi, Jiawei Liu, Hao Tang +1
Vision Language Models (VLMs) excel at identifying and describing objects but often fail at spatial reasoning. We study why VLMs, such as LLaVA, underutilize spatial cues despite h…