activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

Support Operation Factorization: Compositional Readout of Frozen Vision Encoders under Controlled Interventions

Zhongyao Wang, Wanli Ouyang, Taoyong Cui +1

Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separately, however, and can rewar…

cs.CV2026

PolyReal: A Benchmark for Real-World Polymer Science Workflows

Wanhao Liu, Weida Wang, Jiaqing Xie +12

Multimodal Large Language Models (MLLMs) excel in general domains but struggle with complex, real-world science. We posit that polymer science, an interdisciplinary field spanning…

cs.CV2025

A CLIP-Powered Framework for Robust and Generalizable Data Selection

Suorong Yang, Peng Ye, Wanli Ouyang +2

Large-scale datasets have been pivotal to the advancements of deep learning models in recent years, but training on such large datasets invariably incurs substantial storage and co…

cs.CV2025

AutoMat: Enabling Automated Crystal Structure Reconstruction from Microscopy via Agentic Tool Use

Yaotian Yang, Yiwen Tang, Yizhe Chen +14

Reconstructing atomistic crystal structures from a single noisy STEM projection is an ill-posed inverse problem: multiple lattices can explain similar contrast, and purely feed-for…

cs.CV2025

Critic-V: VLM Critics Help Catch VLM Errors in Multimodal Reasoning

Di Zhang, Junxian Li, Jingdi Lei +10

Vision-language models (VLMs) have shown remarkable advancements in multimodal reasoning tasks. However, they still often generate inaccurate or irrelevant responses due to issues…