Publications (6)
An Extensive Benchmark for Single-round and Multi-round Instruction-based Image Editing
Yiwei Ma, Ke Ye, Weihuang Lin +4
In recent years, there have been notable advancements in the area of instruction-based image editing (IIE), which focuses on the automatic alteration of input images using a model.…
I2EBench: A Comprehensive Benchmark for Instruction-based Image Editing
Yiwei Ma, Jiayi Ji, Ke Ye +6
Significant progress has been made in the field of Instruction-based Image Editing (IIE). However, evaluating these models poses a significant challenge. A crucial requirement in t…
VoiceGiraffe: A Benchmark for Extreme Long-Context Audio-Language Understanding
Jashin Ye, Dongxiao Wang, Yixuan Ye +10
While large audio language models (LALMs) have achieved remarkable progress in audio processing at the second- or minute-level scale, understanding hour-level audio remains a funda…
INF-LLaVA: Dual-perspective Perception for High-Resolution Multimodal Large Language Model
Yiwei Ma, Zhibin Wang, Xiaoshuai Sun +4
With advancements in data availability and computing resources, Multimodal Large Language Models (MLLMs) have showcased capabilities across various fields. However, the quadratic c…
CIR-CoT: Towards Interpretable Composed Image Retrieval via End-to-End Chain-of-Thought Reasoning
Weihuang Lin, Yiwei Ma, Jiayi Ji +2
Composed Image Retrieval (CIR), which aims to find a target image from a reference image and a modification text, presents the core challenge of performing unified reasoning across…
HRSeg: High-Resolution Visual Perception and Enhancement for Reasoning Segmentation
Weihuang Lin, Yiwei Ma, Xiaoshuai Sun +4
The reasoning segmentation task involves segmenting objects within an image by interpreting implicit user instructions, which may encompass subtleties such as contextual cues and o…