collaborators

6 papers

cs.CE2026

The Objective Decides: When a Learned Dynamics Model Uses a Conserved Quantity

Chih-Ting Liao, Xin Cao

A linear probe that recovers a conserved quantity from a learned dynamics model's activations is routinely read as evidence that the model uses that quantity. We show this inferenc…

cs.CV2026

Present but Not Remembered: Auditing How Frozen VLAs Encode, Deploy, and Steer Visual History

Chih-Ting Liao, Xin Cao

A frozen vision-language-action model (VLA) receives recent observations at every decision step, yet prior work has focused on adding memory rather than asking how existing history…

cs.CV2026

Brand-as-Memory: Vision-Language Models Encode Causal, Mechanistically Localizable Credibility Priors for News Sources

Chih-Ting Liao, Xin Cao

Vision-language models (VLMs) increasingly read news and web content as images, where the publisher's identity is visually present. We show that VLMs carry a strong source-credibil…

cs.CV2026

Decodable Is Not Grounded: A Vision-Ablation Arbiter for VLM Spatial Reasoning

Chih-Ting Liao, Fei Shen, Xin Cao +1

The standard way to read latent knowledge out of a model, a linear probe confirmed by a steering recovery, can systematically overstate what a vision-language model (VLM) actually…

cs.CV2026

SpaMEM: Benchmarking Dynamic Spatial Reasoning via Perception-Memory Integration in Embodied Environments

Chih-Ting Liao, Xi Xiao, Chunlei Meng +6

Multimodal large language models (MLLMs) have advanced static visual--spatial reasoning, yet they often fail to preserve long-horizon spatial coherence in embodied settings where b…

cs.AI2026

Do LLMs Build World Models From Text? A Multilingual Diagnostic of Spatial Reasoning

Zhikai Pan, Chih-Ting Liao, Chunrui Liu +5

Whether large language models (LLMs) construct internal spatial world models from pure-text descriptions remains contested, and whether such capabilities transfer across languages…