activity
20242026
collaborators

11 papers

cs.AI2026

Evaluating Counterfactual Sensitivity to Patient Information in Medication-Safety Reasoning

Zhitian Hou, Yuhang Liu, Pengkai Wang +8

Applying a valid medication-safety rule when its patient-specific conditions are not met can produce an incorrect decision. Existing medical evaluations largely use isolated and fi…

cs.LG2026

DART: Decoded Attention over Recurrent States for Efficient Long-Context Sequence Modeling

Yixiao Qian, Song Chen, Pengkai Wang +3

Modern language models are built primarily from Transformers, recurrent models, and their hybrid architectures. Transformers rely on token-level attention memories, while recurrent…

cs.CL2026

Beyond Captions: Context-Grounded Reconstruction for Biomedical Multimodal Continued Pretraining

Guanghao Zhu, Zeyu Liu, Zhitian Hou +10

Biomedical figures are explained not by captions alone but by body-text passages that discuss them. Yet current multimodal corpora typically reduce figures to isolated image-captio…

cs.LG2026

Not All Disagreement Is Learnable: Token Teachability in On-Policy Distillation

Yuanyi Wang, Su Lu, Yanggan Gu +6

On-policy distillation (OPD) trains a student on its own rollouts with token-level teacher supervision. Recent selective OPD methods exploit the non-uniformity of OPD signals by pr…

cs.CL2026

E-PMQ: Expert-Guided Post-Merge Quantization with Merged-Weight Anchoring

Wenjun Wang, Yanggan Gu, Shuo Cai +4

Low-resource deployment constraints have made model quantization essential for deploying neural networks while preserving performance. Meanwhile, model merging has become an increa…

cs.LG2026

Discovering Physical Directions in Weight Space: Composing Neural PDE Experts

Pengkai Wang, Pengwei Liu, Yuanyi Wang +7

Recent advances in neural operators have made partial differential equation (PDE) surrogate modeling increasingly scalable and transferable through large-scale pretraining and in-c…