3 papers
cs.LG2026
KForge: LLM-Driven Cross-Platform Kernel Generation for AI Accelerators
Taras Sereda, Burak Bartan, Ankita Nayak +3
Production inference increasingly targets a heterogeneous mix of accelerators. Agentic pipelines interleave reasoning, tool calls, and multi-agent coordination, each with distinct…
cs.LG2025
KForge: Program Synthesis for Diverse AI Hardware Accelerators
Taras Sereda, Tom St. John, Burak Bartan +3
GPU kernels are critical for ML performance but difficult to optimize across diverse accelerators. We present KForge, a platform-agnostic framework built on two collaborative LLM-b…
cs.AR2025
MLPerf Power: Benchmarking the Energy Efficiency of Machine Learning Systems from Microwatts to Megawatts for Sustainable AI
Arya Tschand, Arun Tejusve Raghunath Rajan, Sachin Idgunji +23
Rapid adoption of machine learning (ML) technologies has led to a surge in power consumption across diverse systems, from tiny IoT devices to massive datacenter clusters. Benchmark…