From the 1 of 6 linked papers with an AI index.
6 papers
From Critic to Confidence: PPO for Language-Based Quantitative Prediction with Confidence Estimation
Mehak Dhaliwal, Rasta Tadayon, Andong Hua +2
The paper introduces CARE-PPO, a reinforcement‑learning framework that fine‑tunes large language models to make accurate numeric predictions while simultaneously learning confidenc…
TokenSwap: Benchmarking and Reducing the Modality Gap in Multimodal LLMs
Andong Hua, Colton Bishop, Igor Mordatch +5
Multimodal large language models (MLLMs) should generate consistent responses given semantically equivalent inputs across modalities. However, we observe a systematic discrepancy i…
English is Not All You Need: Systematically Exploring the Role of Multilinguality in LLM Post-Training
Mehak Dhaliwal, Shashwat Chaurasia, Yao Qin +2
Despite the widespread multilingual deployment of large language models, post-training pipelines remain predominantly English-centric, contributing to performance disparities acros…
NutriBench: A Dataset for Evaluating Large Language Models on Nutrition Estimation from Meal Descriptions
Andong Hua, Mehak Preet Dhaliwal, Laya Pullela +2
Accurate nutrition estimation helps people make informed dietary choices and is essential in the prevention of serious health complications. We present NutriBench, the first public…
Flaw or Artifact? Rethinking Prompt Sensitivity in Evaluating LLMs
Andong Hua, Kenan Tang, Chenhe Gu +3
Prompt sensitivity, referring to the phenomenon where paraphrasing (i.e., repeating something written or spoken using different words) leads to significant changes in large languag…
Improving Adversarial Transferability in MLLMs via Dynamic Vision-Language Alignment Attack
Chenhe Gu, Jindong Gu, Andong Hua +1
Multimodal Large Language Models (MLLMs), built upon LLMs, have recently gained attention for their capabilities in image recognition and understanding. However, while MLLMs are vu…