collaborators

11 papers

cs.AI2026

TRACE-ROUTER: Task-Consistent and Adaptive Online Routing for Agentic AI

Ritik Raj, Souvik Kundu, Sarbartha Banerjee +3

Routing to select large language models (LLMs) with different cost-quality trade-offs has become a fundamental deployment feature of enterprise AI. Existing routers, primarily make…

cs.AR2026

SCALE-Sim EVA: Design Principles for an Extensible, Visualizable, and Adaptable Accelerator Simulation Framework

Jingtian Dang, Ritik Raj, Tushar Krishna

Modern AI accelerators increasingly combine heterogeneous compute units, hierarchical memories, local buffers, and specialized data movement paths. This diversity makes fixed accel…

cs.AI2026

Towards Understanding, Analyzing, and Optimizing Agentic AI Execution: A CPU-Centric Perspective

Ritik Raj, Souvik Kundu, Ishita Vohra +2

Agentic AI serving converts monolithic LLM-based inference to autonomous problem-solvers that can plan, call tools, perform reasoning, and adapt on the fly. Due to diverse task exe…

cs.AR2026

SCALE-Sim TPU: Validating and Extending SCALE-Sim for TPUs

Jingtian Dang, Ritik Raj, Changhai Man +2

Cycle-accurate simulators are widely used to study systolic accelerators, yet their accuracy and usability are often limited by weak validation against real hardware and poor integ…

cs.AR2025

OneDSE: Metric-Conditioned Inverse Modeling and Active Search for Sample-Efficient DSE

Ritik Raj, Akshat Ramachandran, Danny Samuel +3

We identify two key challenges in prior CPU design space exploration (DSE) approaches: (a) short-horizon prediction is forward-only: modeling PPA metrics from design parameters whi…

cs.AR2025

Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Abhimanyu Bambhaniya, Ritik Raj, Geonhwa Jeong +6

Large language models (LLMs) have shown remarkable performance across a wide range of applications, often outperforming human experts. However, deploying these gigantic models effi…