5 papers · 1 filter
From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation
Mehak Dhaliwal, Rasta Tadayon, Andong Hua +2
LLMs can perform language-based quantitative prediction from unstructured inputs, but remain susceptible to hallucinations and overconfident errors, making it critical to know not…
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
Andong Hua, Colton Bishop, Igor Mordatch +5
Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy i…
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
Mehak Dhaliwal, Shashwat Chaurasia, Yao Qin +2
Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to performance disparities acros…
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
Andong Hua, Kenan Tang, Chenhe Gu +3
Prompt sensitivity, referring to the phenomenon where paraphrasing (i.e., repeating something written or spoken using different words) leads to significant changes in large languag…
NutriBench: A Dataset for Evaluating Large Language Models on Nutrition Estimation from Meal Descriptions
Andong Hua, Mehak Preet Dhaliwal, Laya Pullela +2
Accurate nutrition estimation helps people make informed dietary choices and is essential in the prevention of serious health complications. We present NutriBench, the first public…