most citedGrounding Long-Context Reasoning with Contextual Normalization for Retrieval-Augmented Generation

1 citations · 1 across the 11 of their papers we have counts for

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

STAMP: Provenance-Guided Credit Assignment for Deep Search Agents

Ke Xu, Han Xu, Xinran Chen +6

Reinforcement learning for deep-search agents has largely focused on trajectory-level scoring -- outcome correctness, citation-aware rewards, and evidence coverage. Yet the actions…

cs.AI2026

Joint Agent Memory and Exploration Learning via Novelty Signals

Shizuo Tian, Xiaohong Weng, Rui Kong +9

In open-ended environments, exploration is fundamental for autonomous agents, yet current language model agents struggle with this. Effective exploration requires memory, but retai…

cs.AI2026

AdaFuse: Accelerating Dynamic Adapter Inference via Token-Level Pre-Gating and Fused Kernel Optimization

Qiyang Li, Rui Kong, Yuchen Li +5

The integration of dynamic, sparse structures like Mixture-of-Experts (MoE) with parameter-efficient adapters (e.g., LoRA) is a powerful technique for enhancing Large Language Mode…

cs.AI2026

Not All Preferences Deserve Gradients: Understanding Gradient Utility in Offline Reasoning Alignment

Hui Wu, Hengyi Cai, Jinman Zhao +6

Offline preference optimization aligns reasoning models from fixed chosen--rejected pairs, yet standard methods apply gradient updates from every pair regardless of its training va…

cs.AI2026

Adversarial Yet Cooperative: Multi-Perspective Reasoning in Retrieved-Augmented Language Models

Can Xu, Lingyong Yan, Jiayi Wu +6

Recent advances in synergizing large reasoning models (LRMs) with retrieval-augmented generation (RAG) have shown promising results, yet two critical challenges remain: (1) reasoni…

cs.AI2025

Efficient Thought Space Exploration Through Strategic Intervention

Ziheng Li, Hengyi Cai, Xiaochi Wei +4

While large language models (LLMs) demonstrate emerging reasoning capabilities, current inference-time expansion methods incur prohibitive computational costs by exhaustive samplin…