5 papers
Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning
Kaizhen Tan, Yang Feng, Heqing Du +3
Metric questions about video require vision-language models to use supplied real-world references to convert visual measurements into physical units. Yet we find that current model…
Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution
Kaizhen Tan, Xin Xu, Siru Tao +4
World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different…
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
Kaizhen Tan, Xin Xu, Siru Tao +4
A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment. Which physical quantities does a train…
Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction
Xin Xu, Chengrui Wu, Jiayu Lu +3
Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered "no" by a conspiracy that is nonetheless profitable…
GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation
Kaizhen Tan, Hanzhe Hong, Siru Tao
Text-to-image models can generate visually plausible city streets, but whether their outputs correspond to a requested road segment rather than a generic city prior remains unclear…