activity
20242026
collaborators

6 papers

cs.RO2026

Beyond Episodic Evaluation: Memory Architectural Bottlenecks in Sequential Embodied Question Answering

Zikui Cai, Kaushal Janga, Tan Dat Dao +15

Embodied question answering (EQA) is traditionally evaluated under an episodic formulation, where agents solve each task independently and reset internal state between episodes. Ho…

cs.CV2026

Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models

Ruchit Rawal, Reza Shirkavand, Sayak Paul +5

Inference-time scaling for text-to-image generation has progressed from simple Best-of- (BoN) sampling to guided search methods that verify and steer candidate trajectories at i…

cs.AI2026

When AI Benchmarks Plateau: A Systematic Study of Benchmark Saturation

Mubashara Akhtar, Anka Reuel, Prajna Soni +36

Artificial intelligence benchmarks are an important mechanism to measure model progress and guide deployment decisions. However, benchmarks quickly "saturate", making it difficult…

cs.SE2025

Benchmarking Correctness and Security in Multi-Turn Code Generation

Ruchit Rawal, Jeffrey Yang Fan Chiang, Chihao Shen +4

AI coding assistants powered by large language models (LLMs) have transformed software development, significantly boosting productivity. While existing benchmarks evaluate the corr…

cs.CV2025

ARGUS: Hallucination and Omission Evaluation in Video-LLMs

Ruchit Rawal, Reza Shirkavand, Heng Huang +2

Video large language models have not yet been widely deployed, largely due to their tendency to hallucinate. Typical benchmarks for Video-LLMs rely simply on multiple-choice questi…

cs.SE2024

Hints Help Finding and Fixing Bugs Differently in Python and Text-based Program Representations

Ruchit Rawal, Victor-Alexandru Pădurean, Sven Apel +2

With the recent advances in AI programming assistants such as GitHub Copilot, programming is not limited to classical programming languages anymore--programming tasks can also be e…