6 papers
The Objective Decides: When a Learned Dynamics Model Uses a Conserved Quantity
Chih-Ting Liao, Xin Cao
A linear probe that recovers a conserved quantity from a learned dynamics model's activations is routinely read as evidence that the model uses that quantity. We show this inferenc…
Present but Not Remembered: Auditing How Frozen VLAs Encode, Deploy, and Steer Visual History
Chih-Ting Liao, Xin Cao
A frozen vision-language-action model (VLA) receives recent observations at every decision step, yet prior work has focused on adding memory rather than asking how existing history…
Brand-as-Memory: Vision-Language Models Encode Causal, Mechanistically Localizable Credibility Priors for News Sources
Chih-Ting Liao, Xin Cao
Vision-language models (VLMs) increasingly read news and web content as images, where the publisher's identity is visually present. We show that VLMs carry a strong source-credibil…
Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning
Chih-Ting Liao, Fei Shen, Xin Cao +1
The standard way to read latent knowledge out of a model, a linear probe confirmed by a steering recovery, can systematically overstate what a vision-language model (VLM) actually…
SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments
Chih-Ting Liao, Xi Xiao, Chunlei Meng +6
Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence in embodied settings where b…
Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning
Zhikai Pan, Chih-Ting Liao, Chunrui Liu +5
Whether large language models (LLMs) construct internal spatial world models from pure-text descriptions remains contested, and whether such capabilities transfer across languages…