5 papers
Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution
Kaizhen Tan, Xin Xu, Siru Tao +4
World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different…
What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations
Kaizhen Tan, Xin Xu, Siru Tao +4
The paper investigates which physical properties (mass, drag, stiffness) are encoded in latent world models by using controlled interventions in a simulated multimodal environment…
Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction
Xin Xu, Chengrui Wu, Jiayu Lu +3
Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered "no" by a conspiracy that is nonetheless profitable…
GeoFidelity-Bench: Evaluating Segment-Level Geographic Fidelity in Text-to-Image Street-View Generation
Kaizhen Tan, Hanzhe Hong, Siru Tao
Text-to-image models can generate visually plausible city streets, but whether their outputs correspond to a requested road segment rather than a generic city prior remains unclear…
When Does Visual Token Pruning Improve Calibration? The Role of Evidence Coverage in MLLMs
Kaizhen Tan, Yang Feng, Heqing Du +3
Visual token pruning is widely used to reduce the inference cost of multimodal large language models (MLLMs), but it is usually evaluated only by accuracy. We study how pruning aff…