Showing cs.MAShow all
2 papers · 1 filter
cs.MA2026
Reuse, Don't Recompute: Efficient Large Reasoning Model Inference via Memory Orchestration
Daivik Patel, Shrenik Patel
Large reasoning models (LRMs) achieve strong accuracy through test-time scaling, generating longer chains of thought or sampling multiple solutions, but at steep costs in tokens an…
cs.MA2026
ENGRAM: Effective, Lightweight Memory Orchestration for Conversational Agents
Daivik Patel, Shrenik Patel
Large language models (LLMs) deployed in user-facing applications require long-horizon consistency: the ability to remember prior interactions, respect user preferences, and ground…