10 papers
Smooth Scaling Laws Hide Stepwise Token Learning
Pingjie Wang, Zechen Hu, Peiru Yang +2
Language model loss follows remarkably regular scaling laws over model and data size, yet it remains unclear why the aggregate loss should exhibit a power-law form. Existing explan…
Mining Useful General Data for Low-Resource Domain Adaptation
Pingjie Wang, Hongcheng Liu, Yusheng Liao +5
Adapting large language models (LLMs) to low-resource domains remains challenging due to the scarcity of domain-specific data. While in-domain data is limited, there exists a vast…
MIRA: Mid-training Rubric Anchoring for Source-Aware Data Selection
Haowen Wang, Yaxin Du, Jian Yang +9
Mid-training has become an important stage in modern LLM development, using large-scale curated mixtures to strengthen capabilities before final post-training. Its data selection p…
Cross-Modal Coreference Alignment: Enabling Reliable Information Transfer in Omni-LLMs
Hongcheng Liu, Yuhao Wang, Zhe Chen +5
Omni Large Language Models (Omni-LLMs) have demonstrated impressive capabilities in holistic multi-modal perception, yet they consistently falter in complex scenarios requiring syn…
When Seeing Is not Enough: Revealing the Limits of Active Reasoning in MLLMs
Hongcheng Liu, Pingjie Wang, Yuhao Wang +3
Multimodal large language models (MLLMs) have shown strong capabilities across a broad range of benchmarks. However, most existing evaluations focus on passive inference, where mod…
Towards Omni-RAG: Comprehensive Retrieval-Augmented Generation for Large Language Models in Medical Applications
Zhe Chen, Yusheng Liao, Shuyang Jiang +4
Large language models hold promise for addressing medical challenges, such as medical diagnosis reasoning, research knowledge acquisition, clinical decision-making, and consumer he…