2 citations · 2 across the 8 of their papers we have counts for
5 papers · 1 filter
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Zichao Lin, Yifeng Xie, Bowen Qu +30
We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmar…
I2E: From Image Pixels to Actionable Interactive Environments for Text-Guided Image Editing
Jinghan Yu, Junhao Xiao, Chenyu Zhu +9
Existing text-guided image editing methods primarily rely on end-to-end pixel-level inpainting paradigm. Despite its success in simple scenarios, this paradigm still significantly…
ARGaze: Autoregressive Transformers for Online Egocentric Gaze Estimation
Jia Li, Wenjie Zhao, Shijian Deng +6
Online egocentric gaze estimation predicts where a camera wearer is looking from first-person video using only past and current frames, a task essential for augmented reality and a…
Toward Gaze Target Detection of Young Autistic Children
Shijian Deng, Erin E. Kosloski, Siva Sai Nagender Vasireddy +6
The automatic detection of gaze targets in autistic children through artificial intelligence can be impactful, especially for those who lack access to a sufficient number of profes…
From Waveforms to Pixels: A Survey on Audio-Visual Segmentation
Jia Li, Yapeng Tian
Audio-Visual Segmentation (AVS) aims to identify and segment sound-producing objects in videos by leveraging both visual and audio modalities. It has emerged as a significant resea…