activity
20242026
collaborators

7 papers

cs.AI2026

CrossCult-KIBench: A Benchmark for Cross-Cultural Knowledge Insertion in MLLMs

Zhen Zeng, Leijiang Gu, Feng Li +2

Multimodal Large Language Models (MLLMs), trained primarily on English-centric data, frequently generate culturally inappropriate or misaligned responses in cross-cultural settings…

cs.CV2026

ProAPO: Progressively Automatic Prompt Optimization for Visual Classification

Xiangyan Qu, Gaopeng Gou, Jiamin Zhuang +5

Vision-language models (VLMs) have made significant progress in image classification by training with large-scale paired image-text data. Their performances largely depend on the p…

cs.CR2025

Respond to Change with Constancy: Instruction-tuning with LLM for Non-I.I.D. Network Traffic Classification

Xinjie Lin, Gang Xiong, Gaopeng Gou +4

Encrypted traffic classification is highly challenging in network security due to the need for extracting robust features from content-agnostic traffic data. Existing approaches fa…

cs.CV2025

Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval

Yuanmin Tang, Jing Yu, Keke Gai +4

Zero-Shot Composed Image Retrieval (ZS-CIR) involves diverse tasks with a broad range of visual content manipulation intent across domain, scene, object, and attribute. The key cha…

cs.CV2025

MADS: Multi-Attribute Document Supervision for Zero-Shot Image Classification

Xiangyan Qu, Jing Yu, Jiamin Zhuang +3

Zero-shot learning (ZSL) aims to train a model on seen classes and recognize unseen classes by knowledge transfer through shared auxiliary information. Recent studies reveal that d…

cs.CV2024

Reason-before-Retrieve: One-Stage Reflective Chain-of-Thoughts for Training-Free Zero-Shot Composed Image Retrieval

Yuanmin Tang, Xiaoting Qin, Jue Zhang +7

Composed Image Retrieval (CIR) aims to retrieve target images that closely resemble a reference image while integrating user-specified textual modifications, thereby capturing user…