6 papers
CHAI: Command Hijacking against embodied AI
Luis Burbano, Diego Ortiz, Qi Sun +5
Embodied Artificial Intelligence (AI) promises to handle edge cases in robotic vehicle systems where data is scarce by using common-sense reasoning grounded in perception and actio…
AHELM: A Holistic Evaluation of Audio-Language Models
Tony Lee, Haoqin Tu, Chi Heem Wong +6
Evaluations of audio-language models (ALMs) -- multimodal models that take interleaved audio and text as input and output text -- are hindered by the lack of standardized benchmark…
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
Yuhan Wang, Siwei Yang, Bingchen Zhao +4
Recent advancements in large multimodal models like GPT-4o have set a new standard for high-fidelity, instruction-guided image editing. However, the proprietary nature of these mod…
AQA-Bench: An Interactive Benchmark for Evaluating LLMs' Sequential Reasoning Ability
Siwei Yang, Bingchen Zhao, Cihang Xie
This paper introduces AQA-Bench, a novel benchmark to assess the sequential reasoning capabilities of large language models (LLMs) in algorithmic contexts, such as depth-first sear…
Robust and Efficient AI-Based Attack Recovery in Autonomous Drones
Diego Ortiz Barbosa, Luis Burbano, Siwei Yang +4
We introduce an autonomous attack recovery architecture to add common sense reasoning to plan a recovery action after an attack is detected. We outline use-cases of our architectur…
: CoT-Like Instruction Generation for Complexity-Controllable Image Editing Benchmark
Siwei Yang, Mude Hui, Bingchen Zhao +3
We introduce , a comprehensive benchmark designed to systematically evaluate instruction-based image editing models across instructions of varying complexity…