works on

From the 1 of 22 linked papers with an AI index.

activity
20242026
collaborators

22 papers

cs.LG2026

Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution

Kaizhen Tan, Xin Xu, Siru Tao +4

World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different…

cs.LG2026

What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

Kaizhen Tan, Xin Xu, Siru Tao +4

The paper investigates which physical properties (mass, drag, stiffness) are encoded in latent world models by using controlled interventions in a simulated multimodal environment…

cs.LG2026

How Many Bits Can an Adapter Write? Measuring the Capacity and Memorization of Parameter-Efficient Fine-Tuning

Kaizhen Tan, Heqing Du, Yang Feng

A LoRA adapter is a few megabytes that almost everyone treats as a skill rather than a record of the data behind it. We put that assumption on a scale. Extending compression-based…

cs.CV2026

CREG: Compass Relational Evidence Graph for Characterizing Directional Structure in VLM Spatial-Reasoning Attribution

Kaizhen Tan, Yang Feng, Heqing Du

Standard attribution heatmaps show where a vision-language model (VLM) focuses, but they do not reveal whether the recovered evidence is organized by the queried spatial relation o…

cs.CL2026

BayLing-Duplex: Native Full-Duplex Speech Dialogue with a Single Autoregressive LLM

Qingkai Fang, Shoutao Guo, Yang Feng

Real-time, full-duplex speech interaction is a key feature of next-generation spoken chatbots, allowing the model to listen and speak at the same time and to handle natural phenome…

cs.CL2026

FreezeEmpath: Efficient Training for Empathetic Spoken Chatbots with Frozen LLMs

Yun Hong, Yan Zhou, Yang Feng

Empathy is essential for fostering natural interactions in spoken dialogue systems, as it enables machines to recognize the emotional tone of human speech and deliver empathetic re…