From the 1 of 20 linked papers with an AI index.
20 papers
PreDiff-LM: Pretrained Discrete Masked Diffusion Language Modeling with Hybrid Attention
Zhengtao Yao, Runhao Li, Xupeng Chen +12
The paper presents PreDiff-LM, a discrete masked diffusion language model that retains causal attention on the prompt while applying bidirectional attention within masked targets,…
SpecAlign: Efficient Specification-Grounded Alignment of Large Language Models via Synthetic Data
Wenjie Wang, Yue Huang, Zhengqing Yuan +6
As large language models (LLMs) are increasingly deployed in real-world applications, alignment is no longer governed by a single universal notion of safety or helpfulness, but ins…
Food4All: An Agentic Framework and Benchmark for Food Resource Navigation with Adaptive User Understanding
Yiyang Li, Weixiang Sun, Tianyi Ma +3
Food assistance referral requires conversational agents to translate underspecified, often noisy help-seeking dialogues into locally valid resource recommendations. We present Food…
Visual Aesthetic Benchmark: Can Frontier Models Judge Beauty?
Yichen Feng, Yuetai Li, Chunjiang Liu +14
Multimodal large language models (MLLMs) are now routinely deployed for visual understanding, generation, and curation. A substantial fraction of these applications require an expl…
NARRA-Gym for Evaluating Interactive Narrative Agents
Yue Huang, Yuchen Ma, Jiayi Ye +14
Interactive narrative tasks require LLMs to sustain a coherent, evolving story while adapting to a user over multiple turns. However, suitable benchmarks for this setting are limit…
CiteAudit: You Cited It, But Did You Read It? A Benchmark for Verifying Scientific References in the LLM Era
Kaiwen Shi, Weixiang Sun, Zheyuan Zhang +3
Scientific research relies on citation integrity, yet large language models (LLMs) have introduced a critical risk: fabricated references that appear plausible but correspond to no…