2 papers
cs.AI2026
GAR: Carbon-Aware Routing for LLM Inference via Constrained Optimization
Disha Sheshanarayana, Rajat Subhra Pal, Manjira Sinha +1
The growing deployment of large language models (LLMs) makes per-request routing essential for balancing response quality and computational cost across heterogeneous model pools. C…
cs.CL2026
Thinking in Latents: Adaptive Anchor Refinement for Implicit Reasoning in LLMs
Disha Sheshanarayana, Rajat Subhra Pal, Manjira Sinha +1
Token-level Chain-of-Thought (CoT) prompting has become a standard way to elicit multi-step reasoning in large language models (LLMs), especially for mathematical word problems. Ho…