3 citations · 5 across the 16 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
AutoSupervision: Closing the Feedback Loop in Scientific Workflows with Grounded Revision Verification
Haobo Li, Eunseo Jung, Wenxiao Zhao +8
Recent advances in large language models (LLMs) have enabled AI systems to assist scientific research and peer review. However, an essential capability for reliable AI-assisted sci…
cs.CL2026
WildTrace: Benchmarking Natural Evidence Trails in Long-Context Reasoning
Zixin Chen, Peng Liu, Haobo Li +7
Answering complex questions over long documents frequently requires integrating evidence that the source itself disperses naturally across distant passages. In an incident report,…
cs.CL2026
VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs
Yurui Xiang, Xingyi Mao, Rui Sheng +7
Large language models (LLMs) show promise in medical diagnosis, but real-world deployment remains challenging due to high-stakes clinical decisions and imperfect reasoning reliabil…