5 papers
Active Learning with Foundation Model Priors: Efficient Learning under Class Imbalance
Jiancheng Zhang, Meiqing Li, Qi Zhang +1
Real-world datasets across image and text domains are often characterized by skewed class distributions and noisy annotations, which jointly degrade model performance, particularly…
Active Testing of Large Language Models via Approximate Neyman Allocation
Zeli Liu, Jiancheng Zhang, Cong Liu +1
Large language models (LLMs) require reliable evaluation from pre-training to test-time scaling, making evaluation a recurring rather than one-off cost. As model scales grow and ta…
Test-Time Matching: Unlocking Compositional Reasoning in Multimodal Models
Yinglun Zhu, Jiancheng Zhang, Fuzhi Tang
Frontier AI models have achieved remarkable progress, yet recent studies suggest they struggle with compositional reasoning, often performing at or below random chance on establish…
Towards Multimodal Active Learning: Efficient Learning with Limited Paired Data
Jiancheng Zhang, Yinglun Zhu
Active learning (AL) is a principled strategy to reduce annotation cost in data-hungry deep learning. However, existing AL algorithms focus almost exclusively on unimodal data, ove…
Mixtraining: A Better Trade-Off Between Compute and Performance
Zexin Li, Jiancheng Zhang, Yufei Li +2
Incorporating self-supervised learning (SSL) before standard supervised learning (SL) has become a widely used strategy to enhance model performance, particularly in data-limited s…