16 papers
CausalMix: Data Mixture as Causal Inference for Language Model Training
Zinan Tang, Yukun Zhang, Shaomian Zheng +6
In Large Language Model (LLM) training, data mixing plays a pivotal role in determining model performance. Recent methods optimize mixture weights via proxy models, but they rely o…
Stellar: Scalable Multimodal Document Retrieval for Natural Language Queries
Yuxiang Guo, Zhonghao Hu, Yuren Mao +5
Multimodal document retrieval--selecting the most relevant multimodal document from a large corpus to answer a natural language query--plays an essential role in Retrieval-Augmente…
On Representation Redundancy in Large-Scale Instruction Tuning Data Selection
Youwei Shu, Shaomian Zheng, Dingnan Jin +5
Data quality is a crucial factor in large language models training. While prior work has shown that models trained on smaller, high-quality datasets can outperform those trained on…
UniGeM: Unifying Data Mixing and Selection via Geometric Exploration and Mining
Changhao Wang, Yunfei Yu, Xinhao Yao +5
The scaling of Large Language Models (LLMs) is increasingly limited by data quality. Most methods handle data mixing and sample selection separately, which can break the structure…
Ostrakon-VL: Towards Domain-Expert MLLM for Food-Service and Retail Stores
Zhiyong Shen, Gongpeng Zhao, Jun Zhou +10
Multimodal Large Language Models (MLLMs) have recently achieved substantial progress in general-purpose perception and reasoning. Nevertheless, their deployment in Food-Service and…
EvolMem: A Cognitive-Driven Benchmark for Multi-Session Dialogue Memory
Ye Shen, Dun Pei, Yiqiu Guo +6
Despite recent advances in understanding and leveraging long-range conversational memory, existing benchmarks still lack systematic evaluation of large language models(LLMs) across…