activity
20242026
most citedDecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning

5 citations · 6 across the 13 of their papers we have counts for

collaborators
Showing cs.CVShow all

8 papers · 1 filter

cs.CV20265 cited

DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning

Chengxuan Qian, Shuo Xing, Shawn Li +2

Multimodal representation learning aims to capture both shared and complementary semantic information across multiple modalities. However, the intrinsic heterogeneity of diverse mo…

cs.CV2026

Agent Banana: High-Fidelity Image Editing with Agentic Thinking and Tooling

Ruijie Ye, Jiayi Zhang, Zhuoxin Liu +10

We study instruction-based image editing under professional workflows and identify three persistent challenges: (i) editors often over-edit, modifying content beyond the user's int…

cs.CV2026

Human-Aligned MLLM Judges for Fine-Grained Image Editing Evaluation: A Benchmark, Framework, and Analysis

Runzhou Liu, Hailey Weingord, Sejal Mittal +18

Evaluating image editing models remains challenging due to the coarse granularity and limited interpretability of traditional metrics, which often fail to capture aspects important…

cs.CV2026

CMOOD: Concept-based Multi-label OOD Detection

Zhendong Liu, Yi Nian, Yuehan Qin +4

How can models effectively detect out-of-distribution (OOD) samples in complex, multi-label settings without extensive retraining? Existing OOD detection methods struggle to captur…

cs.CV2025

Charts Are Not Images: On the Challenges of Scientific Chart Editing

Shawn Li, Ryan Rossi, Sungchul Kim +5

Generative models, such as diffusion and autoregressive approaches, have demonstrated impressive capabilities in editing natural images. However, applying these tools to scientific…

cs.CV2025

Treble Counterfactual VLMs: A Causal Approach to Hallucination

Shawn Li, Jiashu Qu, Yuxiao Zhou +3

Vision-Language Models (VLMs) have advanced multi-modal tasks like image captioning, visual question answering, and reasoning. However, they often generate hallucinated outputs inc…