works on

From the 1 of 12 linked papers with an AI index.

collaborators

12 papers

cs.CV2026

Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA

Zhongkuan Mao, Xianjie Liu, Tianyu Meng +9

The paper proposes a training‑free, single‑pass method that routes intermediate‑layer visual evidence to improve high‑resolution visual question answering without extra image proce…

cs.CL2026

Maximizing Local Entropy Where It Matters: Prefix-Aware Localized LLM Unlearning

Naixin Zhai, Pengyang Shao, Binbin Zheng +4

Machine unlearning aims to forget sensitive knowledge from Large Language Models (LLMs) while maintaining general utility. However, existing approaches typically treat all tokens i…

cs.CV2026

ASTRA: Let Arbitrary Subjects Transform in Video Editing

Fei Shen, Weihao Xu, Rui Yan +4

While existing video editing methods excel with single subjects, they struggle in dense, multi-subject scenes, frequently suffering from attention dilution and mask boundary entang…

cs.CV2026

Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models

Hao Tang, Yu Liu, Shuanglin Yan +3

Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in…

cs.CV2026

Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration

Jinghan He, Junfeng Fang, Feng Xiong +5

Self-play has enabled large language models to autonomously improve through self-generated challenges. However, existing self-play methods for vision-language models rely on passiv…

cs.AI2026

Seeing through the Conflict: Transparent Knowledge Conflict Handling in Retrieval-Augmented Generation

Hua Ye, Siyuan Chen, Ziqi Zhong +4

Large language models (LLMs) equipped with retrieval--the Retrieval-Augmented Generation (RAG) paradigm--should combine their parametric knowledge with external evidence, yet in pr…