8 papers
C-ReD: A Comprehensive Chinese Benchmark for AI-Generated Text Detection Derived from Real-World Prompts
Chenxi Qing, Junxi Wu, Zheng Liu +5
Recently, large language models (LLMs) are capable of generating highly fluent textual content. While they offer significant convenience to humans, they also introduce various risk…
NaviRAG: Towards Active Knowledge Navigation for Retrieval-Augmented Generation
Jihao Dai, Dingjun Wu, Yuxuan Chen +4
Retrieval-augmented generation (RAG) typically relies on a flat retrieval paradigm that maps queries directly to static, isolated text segments. This approach struggles with more c…
PaperScope: A Multi-Modal Multi-Document Benchmark for Agentic Deep Research Across Massive Scientific Papers
Lei Xiong, Huaying Yuan, Zheng Liu +2
Leveraging Multi-modal Large Language Models (MLLMs) to accelerate frontier scientific research is promising, yet how to rigorously evaluate such systems remains unclear. Existing…
HawkBench: Investigating Resilience of RAG Methods on Stratified Information-Seeking Tasks
Hongjin Qian, Zheng Liu, Chao Gao +3
In real-world information-seeking scenarios, users have dynamic and diverse needs, requiring RAG systems to demonstrate adaptable resilience. To comprehensively evaluate the resili…
Memory-enhanced Retrieval Augmentation for Long Video Understanding
Huaying Yuan, Zheng Liu, Minghao Qin +5
Efficient long-video understanding~(LVU) remains a challenging task in computer vision. Current long-context vision-language models~(LVLMs) suffer from information loss due to comp…
Boosting Long-Context Management via Query-Guided Activation Refilling
Hongjin Qian, Zheng Liu, Peitian Zhang +2
Processing long contexts poses a significant challenge for large language models (LLMs) due to their inherent context-window limitations and the computational burden of extensive k…