3 papers
cs.AI2026
When Does Learning to Stop Help? A Cost-Aware Study of Early Exits in Reasoning Models
Zhe Dong, Fang Qin, Manish Shah
Reasoning models spend test-time compute unevenly across instances, and a growing family of early-exit rules -- confidence thresholds, entropy monitors, answer-stability checks, an…
cs.IR2026
Know Before You Fetch: Calibrated Retrieval-Budget Allocation for Retrieval-Augmented Generation
Zhe Dong, Fang Qin, Manish Shah +1
Retrieval-augmented generation (RAG) typically retrieves a fixed number of passages for every query. This is wasteful when the reader already knows the answer, and it can be harmfu…
cs.IR2026
Diagnosing and Mitigating Retrieval Bottlenecks in LLM-Based Cold-Start Recommendation
Zhe Dong, Fang Qin, Manish Shah +1
Large language models (LLMs) are increasingly used as rerankers in recommender systems, with the expectation that semantic understanding will help in cold-start and long-tail regim…