8 papers
SpecBench: Evaluating Specification-Level Reasoning for Software Engineering LLM Agents
Grant Hamblin, Kevin Song, Zhanda Zhu +4
Software engineering (SWE) agents are transitioning from code generation to full software development lifecycle automation. A critical phase in this lifecycle is specification desi…
PlayGen-MoG: Framework for Diverse Multi-Agent Play Generation via Mixture-of-Gaussians Trajectory Prediction
Kevin Song
Multi-agent trajectory generation in team sports requires models that capture both the diversity of possible plays and realistic spatial coordination between players on plays. Stan…
Decoding Defensive Coverage Responsibilities in American Football Using Factorized Attention Based Transformer Models
Kevin Song, Evan Diewald, Ornob Siddiquee +4
Defensive coverage schemes in the National Football League (NFL) represent complex tactical patterns requiring coordinated assignments among defenders who must react dynamically to…
Evaluating Model-Free Policy Optimization in Masked-Action Environments via an Exact Blackjack Oracle
Kevin Song
Infinite-shoe casino blackjack provides a rigorous, exactly verifiable benchmark for discrete stochastic control under dynamically masked actions. Under a fixed Vegas-style ruleset…
LoRAFusion: Efficient LoRA Fine-Tuning for LLMs
Zhanda Zhu, Qidong Su, Yaoyao Ding +3
Low-Rank Adaptation (LoRA) has become the leading Parameter-Efficient Fine-Tuning (PEFT) method for Large Language Models (LLMs), as it significantly reduces GPU memory usage while…
Aegis: Taxonomy and Optimizations for Overcoming Agent-Environment Failures in LLM Agents
Kevin Song, Anand Jayarajan, Yaoyao Ding +4
Large Language Models (LLMs) agents augmented with domain tools promise to autonomously execute complex tasks requiring human-level intelligence, such as customer service and digit…