collaborators

9 papers

cs.CV2026

DICS: Exploring Data Intrinsic Consistency for Visual Instruction Selection

Yuyang Hong, Jinhui Guo, Jiaqi Gu +6

Visual instruction tuning is crucial for advancing the vision-language alignment and instruction-following capabilities of Vision-Language Models (VLMs). However, identifying optim…

cs.AI2026

DiffPDE: Masked Diffusion Language Models as PDE Solver

Wenxuan Guo, Yuyang Hong, Lubin Fan +4

Existing approaches for synthesizing Partial Differential Equation (PDE) solvers predominantly rely on autoregressive models, yet their global left-to-right decoding incurs substan…

cs.CV2026

Audio-Visual Segmentation via Depth-Guided Collaborative Modeling

Zhaojin Fu, Yuyang Hong, Qi Yang +4

Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects in videos by leveraging both visual and au…

cs.CV2026

SeaVIS: Sound-Enhanced Association for Online Audio-Visual Instance Segmentation

Yingjian Zhu, Ying Wang, Yuyang Hong +5

Recently, an audio-visual instance segmentation (AVIS) task has been introduced, aiming to identify, segment and track individual sounding instances in videos. However, prevailing…

cs.CV2026

CC-VQA: Conflict- and Correlation-Aware Method for Mitigating Knowledge Conflict in Knowledge-Based Visual Question Answering

Yuyang Hong, Jiaqi Gu, Yujin Lou +7

Knowledge-based visual question answering (KB-VQA) demonstrates significant potential for handling knowledge-intensive tasks. However, conflicts arise between static parametric kno…

cs.LG2026

Enhanced Graph Transformer with Serialized Graph Tokens

Ruixiang Wang, Yuyang Hong, Shiming Xiang +1

Transformers have demonstrated success in graph learning, particularly for node-level tasks. However, existing methods encounter an information bottleneck when generating graph-lev…