Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Improving Topic Modeling by Distilling Soft Labels from Language Models
Raymond Li, Amirhossein Abaskohi, Chuyuan Li +2
Traditional neural topic models are typically optimized by reconstructing the document's Bag-of-Words (BoW) representations, overlooking contextual information and struggling with…
cs.CL2026
IntentGrasp: A Comprehensive Benchmark for Intent Understanding
Yuwei Yin, Chuyuan Li, Giuseppe Carenini
Accurately understanding the intent behind speech, conversation, and writing is crucial to the development of helpful Large Language Model (LLM) assistants. This paper introduces I…
cs.CL2025
Confabulations from ACL Publications (CAP): A Dataset for Scientific Hallucination Detection
Federica Gamba, Aman Sinha, Timothee Mickus +12
We introduce the CAP (Confabulations from ACL Publications) dataset, a multilingual resource for studying hallucinations in large language models (LLMs) within scientific text gene…