activity
20242026
most citedWake Vision: A Tailored Dataset and Benchmark Suite for TinyML Computer Vision Applications

4 citations · 6 across the 18 of their papers we have counts for

collaborators
Showing cs.ARShow all

7 papers · 1 filter

cs.AR2026

ArchEval: Measuring AI Agents as Computer Architects

Chenyu Wang, Zishen Wan, Jeffrey Ma +8

Computer architecture has long used benchmarks to make progress measurable. LLM agents create a different measurement problem: success is not merely writing code or tuning paramete…

cs.AR2026

AgentDSE: Reasoning-Augmented Architectural Design Space Exploration

Chenyu Wang, Jiahe Caroline Shi, David Kong +4

Traditional architectural design space exploration (DSE) is highly inefficient, typically requiring tens of thousands of simulator evaluations across various optimization methods.…

cs.AR2025

Modeling and Optimizing Performance Bottlenecks for Neuromorphic Accelerators

Jason Yik, Walter Gallego Gomez, Andrew Cheng +8

Neuromorphic accelerators offer promising platforms for machine learning (ML) inference by leveraging event-driven, spatially-expanded architectures that naturally exploit unstruct…

cs.AR2025

QuArch: A Benchmark for Evaluating LLM Reasoning in Computer Architecture

Shvetank Prakash, Andrew Cheng, Arya Tschand +25

The field of computer architecture, which bridges high-level software abstractions and low-level hardware implementations, remains absent from current large language model (LLM) ev…

cs.AR2025

MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AI

Arya Tschand, Arun Tejusve Raghunath Rajan, Sachin Idgunji +23

Rapid adoption of machine learning (ML) technologies has led to a surge in power consumption across diverse systems, from tiny IoT devices to massive datacenter clusters. Benchmark…

cs.AR2025

QuArch: A Question-Answering Dataset for AI Agents in Computer Architecture

Shvetank Prakash, Andrew Cheng, Jason Yik +14

We introduce QuArch, a dataset of 1500 human-validated question-answer pairs designed to evaluate and enhance language models' understanding of computer architecture. The dataset c…