activity
20232026
most citedEvaluating ChatGPT's Information Extraction Capabilities: An Assessment of Performance, Explainability, Calibration, and Faithfulness

59 citations · 99 across the 40 of their papers we have counts for

collaborators
Showing 2025Show all

16 papers · 1 filter

cs.CL2025

Modeling Uncertainty Trends for Timely Retrieval in Dynamic RAG

Bo Li, Tian Tian, Zhenghua Xu +3

Dynamic retrieval-augmented generation (RAG) allows large language models (LLMs) to fetch external knowledge on demand, offering greater adaptability than static RAG. A central cha…

cs.SE2025

Benchmarking and Studying the LLM-based Agent System in End-to-End Software Development

Zhengran Zeng, Yixin Li, Rui Xie +2

The development of LLM-based autonomous agents for end-to-end software development represents a significant paradigm shift in software engineering. However, the scientific evaluati…

cs.CL2025

Decoupling Safety into Orthogonal Subspace: Cost-Efficient and Performance-Preserving Alignment for Large Language Models

Yutao Mou, Xiaoling Zhou, Yuxiao Luo +2

Safety alignment is essential for building trustworthy artificial intelligence, yet it remains challenging to enhance model safety without degrading general performance. Current ap…

cs.CL2025

AutoRed: A Free-form Adversarial Prompt Generation Framework for Automated Red Teaming

Muxi Diao, Yutao Mou, Keqing He +6

The safety of Large Language Models (LLMs) is crucial for the development of trustworthy AI applications. Existing red teaming methods often rely on seed instructions, which limits…

cs.AI2025

Autoformalizer with Tool Feedback

Qi Guo, Jianing Wang, Jianfei Zhang +8

Autoformalization addresses the scarcity of data for Automated Theorem Proving (ATP) by translating mathematical problems from natural language into formal statements. Efforts in r…

cs.AI20251 cited

TrustJudge: Inconsistencies of LLM-as-a-Judge and How to Alleviate Them

Yidong Wang, Yunze Song, Tingyuan Zhu +11

The adoption of Large Language Models (LLMs) as automated evaluators (LLM-as-a-judge) has revealed critical inconsistencies in current evaluation frameworks. We identify two fundam…