2 papers
cs.CL2026
Illusions of the Gold Standard: A Large-scale Analysis of Human Evaluation Protocols for Long-form Text Generation
Katelyn Xiaoying Mei, Yi-Li Hsu, Minjoon Choi +5
Human evaluation plays a critical role in assessing the quality of generated text. However, the reliability and reproducibility of these evaluations depend on transparent and well-…
cs.CL2026
Clarify or Answer: Reinforcement Learning for Agentic VQA with Context Under-specification
Zongwan Cao, Bingbing Wen, Lucy Lu Wang
Real-world visual question answering (VQA) is often context-dependent: an image-question pair may be under-specified, such that the correct answer depends on external information t…