works on

From the 1 of 9 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CVShow all

6 papers · 1 filter

cs.CV2026

Thinking Beyond Videos: Unifying Video Reasoning and Deep Research for Open-World Video Agents

Wenqi Liu, Shijie Ma, Yunxiao Wang +21

Open-world video understanding often requires a model to locate sparse visual evidence and acquire external knowledge that is absent from the video and its parametric memory. While…

cs.CV2026

FiRE: Enhancing MLLMs with Fine-Grained Context Learning for Complex Image Retrieval

Bohan Hou, Haoqiang Lin, Xuemeng Song +4

The paper introduces an automated pipeline to create a fine-grained multimodal dataset and a two-stage fine-tuning strategy that improves multimodal large language models' ability…

cs.CV2026

InterLV-Search: Benchmarking Interleaved Multimodal Agentic Search

Bohan Hou, Jiuning Gu, Jiayan Guo +5

Existing benchmarks for multimodal agentic search evaluate multimodal search and visual browsing, but visual evidence is either confined to the input or treated as an answer endpoi…

cs.CV2025

ImgEdit: A Unified Image Editing Dataset and Benchmark

Yang Ye, Xianyi He, Zongjian Li +5

Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterpa…

cs.CV2025

ORFormer: Occlusion-Robust Transformer for Accurate Facial Landmark Detection

Jui-Che Chiang, Hou-Ning Hu, Bo-Syuan Hou +4

Although facial landmark detection (FLD) has gained significant progress, existing FLD methods still suffer from performance drops on partially non-visible faces, such as faces wit…

cs.CV2024

Pseudo-triplet Guided Few-shot Composed Image Retrieval

Bohan Hou, Haoqiang Lin, Haokun Wen +3

Composed Image Retrieval (CIR) is a challenging task that aims to retrieve the target image with a multimodal query, i.e., a reference image, and its complementary modification tex…