7 papers
Training Needs Trustworthy Worlds: Verified Synthetic Web Environments for Agent Learning
Chenghao Zhang, Canran Xiao, SaiSai Hu +1
Web agents promise to automate complex digital workflows, but their training remains limited by synthetic environments that look plausible while hiding broken links, inconsistent s…
From Solver Feedback to Faithful Plans: Multi-Role Reinforcement Learning for Symbolic Planning
Chenghao Zhang, Yikai Mao, Shanqi Liu +3
Reliable planning requires converting natural-language instructions into executable symbolic specifications, yet large language models remain brittle without costly PDDL annotation…
ConflictScore: Identifying and Measuring How Language Models Handle Conflicting Evidence
Siyi Liu, Aaron Halfaker, Dan Roth +1
Existing metrics for factuality and faithfulness evaluate whether an answer is supported or contradicted by its grounding documents, but they fail to capture when both supporting a…
Conflicts in Texts: Data, Implications and Challenges
Siyi Liu, Dan Roth
As NLP models become increasingly integrated into real-world applications, it becomes clear that there is a need to address the fact that models often rely on and generate conflict…
Learning Human-Perceived Fakeness in AI-Generated Videos via Multimodal LLMs
Xingyu Fu, Siyi Liu, Yinuo Xu +13
Can humans identify AI-generated (fake) videos and provide grounded reasons? While video generation models have advanced rapidly, a critical dimension -- whether humans can detect…
Towards Long Context Hallucination Detection
Siyi Liu, Kishaloy Halder, Zheng Qi +6
Large Language Models (LLMs) have demonstrated remarkable performance across various tasks. However, they are prone to contextual hallucination, generating information that is eith…