3 papers
cs.SE2026
SWE-Together: Evaluating Coding Agents in Interactive User Sessions
Yifan Wu, Zhuokai Zhao, Songlin Li +8
Most coding-agent benchmarks are static: an agent receives a complete task description up front and is judged only by its final code. Real coding assistance is interactive, with us…
cs.CV2026
PushupBench: Your VLM is not good at counting pushups
Shengzhi Li, Jiarun Chen, Karun Sharma +2
Large vision-language models (VLMs) can recognize \textit{what} happens in video but fail to count \textit{how many} times. We introduce \textbf{PushupBench}, 446 long-form clips (…
cs.CL2024
Abstract2Appendix: Academic Reviews Enhance LLM Long-Context Capabilities
Shengzhi Li, Kittipat Kampa, Rongyu Lin +2
Large language models (LLMs) have shown remarkable performance across various tasks, yet their ability to handle long-context reading remains challenging. This study explores the e…