3 papers
cs.AR2026
TLX: Hardware-Native, Evolvable MIMW GPU Compiler for Large-scale Production Environments
Yue Guan, Hongtao Yu, Peng Chen +10
Modern GPUs increasingly rely on specialized hardware units and asynchronous coordination mechanisms, so performance depends on orchestrating data movement, tensor-core computation…
cs.CV2025
TorchTraceAP: A New Benchmark Dataset for Detecting Performance Anti-Patterns in Computer Vision Models
Hanning Chen, Keyu Man, Kevin Zhu +8
Identifying and addressing performance anti-patterns in machine learning (ML) models is critical for efficient training and inference, but it typically demands deep expertise spann…
cs.SE2025
TritonForge: Profiling-Guided Framework for Automated Triton Kernel Optimization
Haonan Li, Keyu Man, Partha Kanuparthy +6
High-performance GPU kernel optimization remains a critical yet labor-intensive task in modern machine learning workloads. Although Triton, a domain-specific language for GPU progr…