Showing cs.SEShow all
2 papers · 1 filter
cs.SE2026
The Correctness Illusion in LLM-Generated GPU Kernels
Dipankar Sarkar
Benchmarks for LLM-generated GPU kernels (KernelBench, TritonBench, GEAK) score correctness through fixed-shape, small-sample allclose-style checks. The number of inputs varies bet…
cs.SE2026
Before the Pull Request: Mining Multi-Agent Coordination
Dipankar Sarkar
Autonomous coding agents now open millions of pull requests, yet large-scale studies find their PRs are produced faster but accepted less often - a coordination and trust gap that…