From the 1 of 9 linked papers with an AI index.
6 papers · 1 filter
Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents
Wenqi Liu, Shijie Ma, Yunxiao Wang +21
Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While…
FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval
Bohan Hou, Haoqiang Lin, Xuemeng Song +4
The paper introduces an automated pipeline to create a fine-grained multimodal dataset and a two-stage fine-tuning strategy that improves multimodal large language models' ability…
InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search
Bohan Hou, Jiuning Gu, Jiayan Guo +5
Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input or treated as an answer endpoi…
ImgEdit: A Unified Image Editing Dataset and Benchmark
Yang Ye, Xianyi He, Zongjian Li +5
Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterpa…
ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark Detection
Jui-Che Chiang, Hou-Ning Hu, Bo-Syuan Hou +4
Although facial landmark detection (FLD) has gained significant progress, existing FLD methods still suffer from performance drops on partially non-visible faces, such as faces wit…
Pseudo-triplet Guided Few-shot Composed Image Retrieval
Bohan Hou, Haoqiang Lin, Haokun Wen +3
Composed Image Retrieval (CIR) is a challenging task that aims to retrieve the target image with a multimodal query, i.e., a reference image, and its complementary modification tex…