collaborators

14 papers

cs.CV2026

HiST: A Hierarchical Sparse Transformer for Cross-Modal Spatial Transcriptomics Modeling

Weiyi Wu, Xinwen Xu, Xingjian Diao +4

Spatial transcriptomics (ST) links gene expression with tissue morphology but remains expensive and low-throughput, motivating surrogates that infer expression from routine histolo…

cs.CL2026

Doc-to-Atom: Learning to Compile and Compose Memory Atoms

Xingjian Diao, Wenbo Li, Yashas Malur Saidutta +3

Long input sequences are central to document understanding and multi-step reasoning in Large Language Models, yet the quadratic cost of attention makes inference both memory-intens…

cs.CV2026

Addressing Overthinking in Large Vision-Language Models via Gated Perception-Reasoning Optimization

Xingjian Diao, Zheyuan Liu, Chunhui Zhang +6

Large Vision-Language Models (LVLMs) have exhibited strong reasoning capabilities through chain-of-thought mechanisms that generate step-by-step rationales. However, such slow-thin…

cs.SD2026

Music Audio-Visual Question Answering Requires Specialized Multimodal Designs

Wenhao You, Xingjian Diao, Wenjun Huang +9

While recent Multimodal Large Language Models exhibit impressive capabilities for general multimodal tasks, specialized domains like music necessitate tailored approaches. Music Au…

cs.CV2026

Learning Spatial-Preserving Hierarchical Representations for Digital Pathology

Weiyi Wu, Xingjian Diao, Chunhui Zhang +4

Whole slide images (WSIs) pose fundamental computational challenges due to their gigapixel resolution and the sparse distribution of informative regions. Existing approaches often…

cs.CV2026

Exploiting Label-Independent Regularization from Spatial Dependencies for Whole Slide Image Analysis

Weiyi Wu, Xinwen Xu, Chongyang Gao +3

Whole slide images, with their gigapixel-scale panoramas of tissue samples, are pivotal for precise disease diagnosis. However, their analysis is hindered by immense data size and…