4 papers
Explore Briefly, Then Decide: Mitigating LLM Overthinking via Cumulative Entropy Regulation
Yi Bin, Tianyi Jiang, Yujuan Ding +5
Large Language Models (LLMs) have demonstrated remarkable reasoning abilities on complex problems using long Chain-of-Thought (CoT) reasoning. However, they often suffer from overt…
SearchAttack: Red-Teaming LLMs against Knowledge-to-Action Threats under Online Web Search
Yu Yan, Sheng Sun, Mingfeng Li +6
Recently, people have suffered from LLM hallucination and have become increasingly aware of the reliability gap of LLMs in open and knowledge-intensive tasks. As a result, they hav…
Chinese Labor Law Large Language Model Benchmark
Zixun Lan, Maochun Xu, Yifan Ren +7
Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose mod…
D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies
Sen Chen, Tong Zhao, Yi Bin +3
Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial Ge…