activity
20242026
collaborators

6 papers

cs.CV2026

InSight-doc: Agentic Visual Perception for Long-Document Understanding

Kaican Li, Weiyan Xie, Lewei Yao +4

Long-document understanding often requires reasoning over many visually rich pages, making inference costly and prone to context rot. In this work, we propose InSight-doc, an agent…

cs.CV2026

Beyond VLM-Based Rewards: Diffusion-Native Latent Reward Modeling

Gongye Liu, Bo Yang, Yida Zhi +8

Preference optimization for diffusion and flow-matching models relies on reward functions that are both discriminatively robust and computationally efficient. Vision-Language Model…

cs.DC2026

vLLM-Omni: Fully Disaggregated Serving for Any-to-Any Multimodal Models

Peiqi Yin, Jiangyun Zhu, Han Gao +13

Any-to-any multimodal models that jointly handle text, images, video, and audio represent a significant advance in multimodal AI. However, their complex architectures (typically co…

cs.CV2025

CannyEdit: Selective Canny Control and Dual-Prompt Guidance for Training-Free Image Editing

Weiyan Xie, Han Gao, Didan Deng +4

Recent advances in text-to-image (T2I) models have enabled training-free regional image editing by leveraging the generative priors of foundation models. However, existing methods…

cs.AI2025

Reasoning Scaffolding: Distilling the Flow of Thought from LLMs

Xiangyu Wen, Junhua Huang, Zeju Li +6

The prevailing approach to distilling reasoning from Large Language Models (LLMs)-behavioral cloning from textual rationales-is fundamentally limited. It teaches Small Language Mod…

cs.LG2024

Dual Risk Minimization: Towards Next-Level Robustness in Fine-tuning Zero-Shot Models

Kaican Li, Weiyan Xie, Yongxiang Huang +5

Fine-tuning foundation models often compromises their robustness to distribution shifts. To remedy this, most robust fine-tuning methods aim to preserve the pre-trained features. H…