4 papers · 1 filter
SCOUT: Active Information Foraging for Long-Text Understanding with Decoupled Epistemic States
Zhenliang Zhang, Wenqing Wang, Yong Hu +4
Long-Text Understanding (LTU) at million-token scale requires balancing reasoning fidelity with computational efficiency. Frontier long-context LLMs can process millions of token c…
SagaScale: A Realistic, Scalable, and High-Quality Long-Context Benchmark Built from Full-Length Novels
Guancheng Du, Yong Hu, Wenqing Wang +2
Large Language Models (LLMs) have shown significant progress, but understanding long and complex documents remains challenging. Many long-context benchmarks have been proposed, but…
A-IPO: Adaptive Intent-driven Preference Optimization
Wenqing Wang, Muhammad Asif Ali, Ali Shoker +4
Human preferences are diverse and dynamic, shaped by regional, cultural, and social factors. Existing alignment methods like Direct Preference Optimization (DPO) and its variants o…
Defining and Detecting Vulnerability in Human Evaluation Guidelines: A Preliminary Study Towards Reliable NLG Evaluation
Jie Ruan, Wenqing Wang, Xiaojun Wan
Human evaluation serves as the gold standard for assessing the quality of Natural Language Generation (NLG) systems. Nevertheless, the evaluation guideline, as a pivotal element en…