activity
20242026
collaborators

5 papers

stat.ML2026

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

Zilong Zhang, Yi-Ting Hung, Weiyi He +3

Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their prefere…

stat.ML2026

Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning

Zilong Zhang, Yi-Ting Hung, Lei Ding +1

Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic…

cs.LG2025

An Efficient Rehearsal Scheme for Catastrophic Forgetting Mitigation during Multi-stage Fine-tuning

Andrew Bai, Chih-Kuan Yeh, Cho-Jui Hsieh +1

Incrementally fine-tuning foundational models on new tasks or domains is now the de facto approach in NLP. A known pitfall of this approach is the \emph{catastrophic forgetting} of…

cs.CL2024

Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Gemini Team, Petko Georgiev, Ving Ian Lei +1132

In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…

cs.CL2024

Scalable Multi-Domain Adaptation of Language Models using Modular Experts

Peter Schafhalter, Shun Liao, Yanqi Zhou +3

Domain-specific adaptation is critical to maximizing the performance of pre-trained language models (PLMs) on one or multiple targeted tasks, especially under resource-constrained…