6 papers
Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning
Kaizhen Tan, Yang Feng, Heqing Du +3
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units. Yet we find that current model…
You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change
Kaizhen Tan
Vision-language models are increasingly used to measure urban change from repeated street-level imagery, but their longitudinal reliability is not well understood. We test how much…
Vision Models Predict Urban Scene Appraisal with Limited Neural Alignment
Kaizhen Tan, Yuantao Deng
Pretrained vision embeddings are increasingly used as general-purpose representations for modelling how people appraise urban scenes, and are validated almost entirely by how well…
Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution
Kaizhen Tan, Xin Xu, Siru Tao +4
World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different…
Renaming or Tightness: Enforcing Disjunctive Information Flow Policies
Xin Xu, Siru Tao, Kaizhen Tan
A disjunctive policy allows a value to depend on at most one of two secrets and never on both: an analyst may consult one client's file or the other's, a share of a split secret ma…
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
Kaizhen Tan, Xin Xu, Siru Tao +4
A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a train…