collaborators

10 papers

cs.CL2026

Smooth Scaling Laws Hide Stepwise Token Learning

Pingjie Wang, Zechen Hu, Peiru Yang +2

Language model loss follows remarkably regular scaling laws over model and data size, yet it remains unclear why the aggregate loss should exhibit a power-law form. Existing explan…

cs.CL2026

Mining Useful General Data for Low-Resource Domain Adaptation

Pingjie Wang, Hongcheng Liu, Yusheng Liao +5

Adapting large language models (LLMs) to low-resource domains remains challenging due to the scarcity of domain-specific data. While in-domain data is limited, there exists a vast…

cs.AI2026

MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection

Haowen Wang, Yaxin Du, Jian Yang +9

Mid-training has become an important stage in modern LLM development, using large-scale curated mixtures to strengthen capabilities before final post-training. Its data selection p…

cs.CL2026

Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs

Hongcheng Liu, Yuhao Wang, Zhe Chen +5

Omni Large Language Models (Omni-LLMs) have demonstrated impressive capabilities in holistic multi-modal perception, yet they consistently falter in complex scenarios requiring syn…

cs.CL2025

When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs

Hongcheng Liu, Pingjie Wang, Yuhao Wang +3

Multimodal large language models (MLLMs) have shown strong capabilities across a broad range of benchmarks. However, most existing evaluations focus on passive inference, where mod…

cs.CL2025

Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications

Zhe Chen, Yusheng Liao, Shuyang Jiang +4

Large language models hold promise for addressing medical challenges, such as medical diagnosis reasoning, research knowledge acquisition, clinical decision-making, and consumer he…