3 papers
cs.LG2026
Lynx: Enabling Efficient MoE Inference through Dynamic Batch-Aware Expert Selection
Vima Gupta, Jae Hyung Ju, Kartik Sinha +2
Selective parameter activation provided by Mixture-of-Expert (MoE) models have made them a popular choice in modern foundational models. However, MoEs face a fundamental tension wh…
cs.CV2026
CHAI: CacHe Attention Inference for text2video
Joel Mathew Cherian, Ashutosh Muralidhara Bharadwaj, Vima Gupta +1
Text-to-video diffusion models deliver impressive results but remain slow because of the sequential denoising of 3D latents. Existing approaches to speed up inference either requir…
cs.SE2026
SAFuzz: Semantic-Guided Adaptive Fuzzing for LLM-Generated Code
Ziyi Yang, Kalit Inani, Keshav Kabra +2
While AI-coding assistants accelerate software development, current testing frameworks struggle to keep pace with the resulting volume of AI-generated code. Traditional fuzzing tec…