3 citations · 3 across the 4 of their papers we have counts for
10 papers · 1 filter
Thinking Once Is Enough: Intermediate-Layer Evidence Routing for High-Resolution VQA
Zhongkuan Mao, Xianjie Liu, Tianyu Meng +9
High-resolution visual question answering (HR-VQA) is often treated as a problem of insufficient evidence acquisition, where failing multimodal large language models must inspect i…
Active Zero: Self-Evolving Vision-Language Models through Active Environment Exploration
Jinghan He, Junfeng Fang, Feng Xiong +5
Self-play has enabled large language models to autonomously improve through self-generated challenges. However, existing self-play methods for vision-language models rely on passiv…
Cross-modal Proxy Evolving for OOD Detection with Vision-Language Models
Hao Tang, Yu Liu, Shuanglin Yan +3
Reliable zero-shot detection of out-of-distribution (OOD) inputs is critical for deploying vision-language models in open-world settings. However, the lack of labeled negatives in…
SceneDecorator: Towards Scene-Oriented Story Generation with Scene Planning and Scene Consistency
Quanjian Song, Donghao Zhou, Jingyu Lin +5
Recent text-to-image models have revolutionized image generation, but they still struggle with maintaining concept consistency across generated images. While existing works focus o…
ASTRA: Let Arbitrary Subjects Transform in Video Editing
Fei Shen, Weihao Xu, Rui Yan +4
While existing video editing methods excel with single subjects, they struggle in dense, multi-subject scenes, frequently suffering from attention dilution and mask boundary entang…
DictAS: A Framework for Class-Generalizable Few-Shot Anomaly Segmentation via Dictionary Lookup
Zhen Qu, Xian Tao, Xinyi Gong +7
Recent vision-language models (e.g., CLIP) have demonstrated remarkable class-generalizable ability to unseen classes in few-shot anomaly segmentation (FSAS), leveraging supervised…