11 papers
TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI
Ritik Raj, Souvik Kundu, Sarbartha Banerjee +3
Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make…
SCALE-Sim EVA: Design Principles for an Extensible, Visualizable, and Adaptable Accelerator Simulation Framework
Jingtian Dang, Ritik Raj, Tushar Krishna
Modern AI accelerators increasingly combine heterogeneous compute units, hierarchical memories, local buffers, and specialized data movement paths. This diversity makes fixed accel…
Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective
Ritik Raj, Souvik Kundu, Ishita Vohra +2
Agentic AI serving converts monolithic LLM-based inference to autonomous problem-solvers that can plan, call tools, perform reasoning, and adapt on the fly. Due to diverse task exe…
SCALE-Sim TPU: Validating and Extending SCALE-Sim for TPUs
Jingtian Dang, Ritik Raj, Changhai Man +2
Cycle-accurate simulators are widely used to study systolic accelerators, yet their accuracy and usability are often limited by weak validation against real hardware and poor integ…
OneDSE: Metric-Conditioned Inverse Modeling and Active Search for Sample-Efficient DSE
Ritik Raj, Akshat Ramachandran, Danny Samuel +3
We identify two key challenges in prior CPU design space exploration (DSE) approaches: (a) short-horizon prediction is forward-only: modeling PPA metrics from design parameters whi…
Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models
Abhimanyu Bambhaniya, Ritik Raj, Geonhwa Jeong +6
Large language models (LLMs) have shown remarkable performance across a wide range of applications, often outperforming human experts. However, deploying these gigantic models effi…