activity
20242026
collaborators

5 papers

cs.AI2026

ChartAnno: Evaluating MLLMs for Chart Annotation Generation

Zhenghan Chen, Zekai Shao, Lidan Tan +10

Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains under…

cs.AI2026

No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning

Zhicong Li, Lingjie Jiang, Yulan Hu +7

Critique-guided reinforcement learning (RL) has emerged as a powerful paradigm for training LLM agents by augmenting sparse outcome rewards with natural-language feedback. However,…

cs.CV2025

AKRMap: Adaptive Kernel Regression for Trustworthy Visualization of Cross-Modal Embeddings

Yilin Ye, Junchao Huang, Xingchen Zeng +2

Cross-modal embeddings form the foundation for multi-modal models. However, visualization methods for interpreting cross-modal embeddings have been primarily confined to traditiona…

cs.HC2025

GenColor: Generative Color-Concept Association in Visual Design

Yihan Hou, Xingchen Zeng, Yusong Wang +3

Existing approaches for color-concept association typically rely on query-based image referencing, and color extraction from image references. However, these approaches are effecti…

cs.CV2024

ModalChorus: Visual Probing and Alignment of Multi-modal Embeddings via Modal Fusion Map

Yilin Ye, Shishi Xiao, Xingchen Zeng +1

Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable…