activity
20242026
collaborators

6 papers

cs.CV2026

Towards One-to-Many Temporal Grounding

Qi Xu, Yue Tan, Shihao Chen +5

Temporal Grounding (TG) aims to localize video segments corresponding to a textual query. Prior research predominantly focuses on single-segment retrieval. Real-world scenarios, ho…

cs.CV2026

DragNeXt: Rethinking Drag-Based Image Editing

Yuan Zhou, Junbao Zhou, Qingshan Xu +5

Drag-Based Image Editing (DBIE), which allows users to manipulate images by directly dragging objects within them, has recently attracted much attention from the community. However…

cs.CV2025

On Path to Multimodal Generalist: General-Level and General-Bench

Hao Fei, Yuan Zhou, Juncheng Li +29

The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of LLMs. Unlike earlier specialists, existing MLLMs are evolv…

cs.CV2025

Iris: Breaking GUI Complexity with Adaptive Focus and Self-Refining

Zhiqi Ge, Juncheng Li, Xinglei Pang +7

Digital agents are increasingly employed to automate tasks in interactive digital environments such as web pages, software applications, and operating systems. While text-based age…

cs.CV2024

Unified Generative and Discriminative Training for Multi-modal Large Language Models

Wei Chow, Juncheng Li, Qifan Yu +7

In recent times, Vision-Language Models (VLMs) have been trained under two predominant paradigms. Generative training has enabled Multimodal Large Language Models (MLLMs) to tackle…

cs.CV2024

Towards Unified Multimodal Editing with Enhanced Knowledge Collaboration

Kaihang Pan, Zhaoyu Fan, Juncheng Li +6

The swift advancement in Multimodal LLMs (MLLMs) also presents significant challenges for effective knowledge editing. Current methods, including intrinsic knowledge editing and ex…