1 paper · 1 filter
Ammar Ahmed, Azal Ahmad Khan, Ayaan Ahmad +3
Large reasoning models improve accuracy by producing long reasoning traces, but this inflates latency and cost, motivating inference-time efficiency. We propose Retrieval-of-Though…