collaborators

27 papers

cs.CV2026

Technical Report on the CVPR 2026@AdvML Workshop Challenge

Tianyuan Zhang, Zonglei Jing, Jiangfan Liu +47

The paper reports on the CVPR 2026@AdvML Workshop Challenge, which evaluated adversarial attacks on multimodal vision‑language agents for autonomous driving using multi‑view visual…

cs.CV2026

RATS! Patches Talk Through Registers: Emergent Parts in Register Attention Transformers

Timing Yang, Predrag Neskovic, Jansen Seheult +4

When humans see a bird, they recognize far more than just "bird" -- they see a head, wings, and talons, a structured assembly of reusable parts that can be identified across every…

cs.CV2026

SEMAGIC: Learning Semantically Consistent Deformable 3D Representations from In-the-Wild Images

Sky Cen, Wufei Ma, Guofeng Zhang +2

Learning deformable 3D object models from single-view in-the-wild images has enabled impressive 3D shape reconstruction without supervision. However, it remains unclear whether the…

cs.CV2026

Can These Views Be One Scene? Evaluating Multiview 3D Consistency when 3D Foundation Models Hallucinate

Soumava Paul, Prakhar Kaushik, Alan Yuille

Multiview 3D evaluation assumes that the images being scored are observations of one static 3D scene. This assumption can fail in NVS and sparse-view reconstruction: inputs or gene…

cs.AI2026

Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning

Qihao Liu, Luoxin Ye, Wufei Ma +2

Large language models (LLMs) with explicit reasoning capabilities excel at mathematical reasoning yet still commit process errors, such as incorrect calculations, brittle logic, an…

cs.AI2026

Meissa: Multi-modal Medical Agentic Intelligence

Yixiong Chen, Xinyi Bai, Yue Pan +2

Multi-modal large language models (MM-LLMs) have shown strong performance in medical image understanding and clinical reasoning. Recent medical agent systems extend them with tool…