25 papers
CyberForge: Verified Vulnerability Injection at Repository Level for Cybersecurity Agent Training
Amine Lbath, Manan Suri, Aurelien Delaitre +4
Despite recent advances, frontier large language model (LLM) agents remain limited in discovering and patching complex vulnerabilities in real-world software. Generally available a…
Personalized Embodied Navigation for Portable Object Finding
Vishnu Sashank Dorbala, Bhrij Patel, Amrit Singh Bedi +1
Embodied navigation methods commonly operate in static environments with stationary objects. In this work, we present approaches for tackling navigation in dynamic scenarios with n…
Code Comprehension then Auditing for Unsupervised LLM Evaluation
Bhrij Patel, Souradip Chakraborty, Mengdi Wang +2
Large Language Models (LLMs) for unsupervised code correctness evaluation have recently gained attention because they can judge if code runs as intended without requiring reference…
Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning
Chao-Han Huck Yang, Sreyan Ghosh, Qing Wang +14
We present Task 5 of the DCASE 2025 Challenge: an Audio Question Answering (AQA) benchmark spanning multiple domains of sound understanding. This task defines three QA subsets (Bio…
LoLep: Single-View View Synthesis with Locally-Learned Planes and Self-Attention Occlusion Inference
Cong Wang, Yu-Ping Wang, Dinesh Manocha
We propose a novel method, LoLep, which regresses Locally-Learned planes from a single RGB image to represent scenes accurately, thus generating better novel views. Without the dep…
TAC: Timestamped Audio Captioning
Sonal Kumar, Prem Seetharaman, Ke Chen +8
Large Audio Language Models struggle to disentangle overlapping events in complex acoustic scenes, yielding temporally inconsistent captions and frequent hallucinations. We introdu…