19 citations · 48 across the 5 of their papers we have counts for
4 papers · 1 filter
OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
Pengfei Zhou, Xiaopeng Peng, Jiajun Song +15
Multimodal Large Language Models (MLLMs) have made significant strides in visual understanding and generation tasks. However, generating interleaved image-text content remains a ch…
Synthesizing Knowledge-enhanced Features for Real-world Zero-shot Food Detection
Pengfei Zhou, Weiqing Min, Jiajun Song +2
Food computing brings various perspectives to computer vision like vision-based food analysis for nutrition and health. As a fundamental task in food computing, food detection need…
SeeDS: Semantic Separable Diffusion Synthesizer for Zero-shot Food Detection
Pengfei Zhou, Weiqing Min, Yang Zhang +3
Food detection is becoming a fundamental task in food computing that supports various multimedia applications, including food recommendation and dietary monitoring. To deal with re…
ISDA: Position-Aware Instance Segmentation with Deformable Attention
Kaining Ying, Zhenhua Wang, Cong Bai +1
Most instance segmentation models are not end-to-end trainable due to either the incorporation of proposal estimation (RPN) as a pre-processing or non-maximum suppression (NMS) as…