4 papers
AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing
Zilong Zhang, Yi-Ting Hung, Weiyi He +3
Large language models (LLMs) are increasingly used as judges for open-ended generation, as large-scale human evaluation is often expensive and difficult to scale, yet their prefere…
Quantifying and Auditing LLM Evaluation via Positive--Unlabeled Learning
Zilong Zhang, Yi-Ting Hung, Lei Ding +1
Large Language Models (LLMs) are increasingly used as judges for scalable evaluation, yet such LLM--as--a--Judge systems exhibit systematic biases that are decoupled from semantic…
An Efficient Rehearsal Scheme for Catastrophic Forgetting Mitigation during Multi-stage Fine-tuning
Andrew Bai, Chih-Kuan Yeh, Cho-Jui Hsieh +1
Incrementally fine-tuning foundational models on new tasks or domains is now the de facto approach in NLP. A known pitfall of this approach is the \emph{catastrophic forgetting} of…
Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Gemini Team, Petko Georgiev, Ving Ian Lei +1132
In this report, we introduce the Gemini 1.5 family of models, representing the next generation of highly compute-efficient multimodal models capable of recalling and reasoning over…