28 citations · 28 across the 7 of their papers we have counts for
5 papers · 1 filter
CurveShift: Is Agent Progress Scalar? Separating Level from Shape
Hanwen Xing, Pengyun Wang, BingXu Meng +8
Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summa…
Disparities In Negation Understanding Across Languages In Vision-Language Models
Charikleia Moraitaki, Sarah Pan, Skyler Pulling +3
Vision-language models (VLMs) exhibit affirmation bias: a systematic tendency to select positive captions ("X is present") even when the correct description contains negation ("no…
Medical Hallucinations in Foundation Models and Their Impact on Healthcare
Yubin Kim, Hyewon Jeong, Shan Chen +24
Hallucinations in foundation models arise from autoregressive training objectives that prioritize token-likelihood optimization over epistemic accuracy, fostering overconfidence an…
Train Long, Think Short: Curriculum Learning for Efficient Reasoning
Hasan Abed Al Kader Hammoud, Kumail Alhamoud, Abed Hammoud +3
Recent work on enhancing the reasoning abilities of large language models (LLMs) has introduced explicit length control as a means of constraining computational cost while preservi…
MedPAIR: Measuring Physicians and AI Relevance Alignment in Medical Question Answering
Yuexing Hao, Kumail Alhamoud, Hyewon Jeong +6
Large Language Models (LLMs) have demonstrated remarkable performance on various medical question-answering (QA) benchmarks, including standardized medical exams. However, correct…