10 papers
A Latent Variable Framework for Scaling Laws in Large Language Models
Peiyao Cai, Chengyu Cui, Felipe Maia Polo +6
We propose a statistical framework built on latent variable modeling for scaling laws of large language models (LLMs). Our work is motivated by the rapid emergence of numerous new…
Rich Insights from Cheap Signals: Efficient Evaluations via Tensor Factorization
Felipe Maia Polo, Aida Nematzadeh, Virginia Aglietti +2
Moving beyond evaluations that collapse performance across heterogeneous prompts toward fine-grained evaluation at the prompt level, or within relatively homogeneous subsets, is ne…
Bridging Human and LLM Judgments: Understanding and Narrowing the Gap
Felipe Maia Polo, Xinhe Wang, Mikhail Yurochkin +3
Large language models are increasingly used as judges (LLM-as-a-judge) to evaluate model outputs at scale, but their assessments often diverge systematically from human judgments.…
Sloth: scaling laws for LLM skills to predict multi-benchmark performance across families
Felipe Maia Polo, Seamus Somerstep, Leshem Choshen +2
Scaling laws for large language models (LLMs) predict model performance based on parameters like size and training data. However, differences in training configurations and data pr…
Personalized Image Generation for Recommendations Beyond Catalogs
Gabriel Patron, Zhiwei Xu, Ishan Kapnadak +1
Personalization is central to human-AI interaction, yet current diffusion-based image generation systems remain largely insensitive to user diversity. Existing attempts to address…
COMET-poly: Machine Translation Metric Grounded in Other Candidates
Maike Züfle, Vilém Zouhar, Tu Anh Dinh +3
Automated metrics for machine translation attempt to replicate human judgment. Unlike humans, who often assess a translation in the context of multiple alternatives, these metrics…