activity
20232026
collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation

Mehak Dhaliwal, Rasta Tadayon, Andong Hua +2

LLMs can perform language-based quantitative prediction from unstructured inputs, but remain susceptible to hallucinations and overconfident errors, making it critical to know not…

cs.CL2026

TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs

Andong Hua, Colton Bishop, Igor Mordatch +5

Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy i…

cs.CL2026

English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training

Mehak Dhaliwal, Shashwat Chaurasia, Yao Qin +2

Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to performance disparities acros…

cs.CL2025

Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs

Andong Hua, Kenan Tang, Chenhe Gu +3

Prompt sensitivity, referring to the phenomenon where paraphrasing (i.e., repeating something written or spoken using different words) leads to significant changes in large languag…

cs.CL2024

NutriBench: A Dataset for Evaluating Large Language Models on Nutrition Estimation from Meal Descriptions

Andong Hua, Mehak Preet Dhaliwal, Laya Pullela +2

Accurate nutrition estimation helps people make informed dietary choices and is essential in the prevention of serious health complications. We present NutriBench, the first public…