collaborators

7 papers

cs.CV2026

Open-Ended CT Volume Segmentation with Weak Supervision from Language

Sanjay Subramanian, Junwei Yu, Zirui Wang +5

We introduce a method for training a text-conditioned segmentation model for CT scans, which combines voxel-level supervision with coarse but scalable slice-level supervision from…

cs.CV2026

Stateful Visual Encoders for Vision-Language Models

Zirui Wang, Junwei Yu, Adam Yala +3

Vision-language models (VLMs) are increasingly used in multi-image, multi-turn agentic settings where decisions depend on visual changes. However, in existing open-weight VLMs, vis…

cs.IR2026

PIXELRAG: Web Screenshots Beat Text for Retrieval-Augmented Generation

Yichuan Wang, Zhifei Li, Zirui Wang +5

Augmenting large language models (LLMs) with retrieved web text has become a dominant paradigm, yet the web is not natively textual: existing systems depend on complex parsing pipe…

cs.CV2026

AEGIS: A Holistic Benchmark for Evaluating Forensic Analysis of AI-Generated Academic Images

Bo Zhang, Tzu-Yen Ma, Zichen Tang +18

We introduce AEGIS, A holistic benchmark for Evaluating forensic analysis of AI-Generated academic ImageS. Compared to existing benchmarks, AEGIS features three key advances: (1) D…

cs.CV2026

THEMIS: Towards Holistic Evaluation of MLLMs for Scientific Paper Fraud Forensics

Tzu-Yen Ma, Bo Zhang, Zichen Tang +20

We present THEMIS, a novel multi-task benchmark designed to comprehensively evaluate multimodal large language models (MLLMs) on visual fraud reasoning within real-world academic s…

cs.CV2026

Kestrel: Grounding Self-Refinement for LVLM Hallucination Mitigation

Jiawei Mao, Hardy Chen, Haoqin Tu +7

Large vision-language models (LVLMs) have become increasingly strong but remain prone to hallucinations in multimodal tasks, which significantly narrows their deployment. As traini…