Publications (43)
Polyp-Gen: Realistic and Diverse Polyp Image Generation for Endoscopic Dataset Expansion
Shengyuan Liu, Zhen Chen, Qiushi Yang +4
Automated diagnostic systems (ADS) have shown significant potential in the early detection of polyps during endoscopic examinations, thereby reducing the incidence of colorectal ca…
EndoSparse: Real-Time Sparse View Synthesis of Endoscopic Scenes using Gaussian Splatting
Chenxin Li, Brandon Y. Feng, Yifan Liu +4
3D reconstruction of biological tissues from a collection of endoscopic images is a key to unlock various important downstream surgical applications with 3D capabilities. Existing…
Artificial Hippocampus Networks for Efficient Long-Context Modeling
Yunhao Fang, Weihao Yu, Shu Zhong +3
Long-sequence modeling faces a fundamental trade-off between the efficiency of compressive fixed-size memory in RNN-like models and the fidelity of lossless growing memory in atten…
Mugs: A Multi-Granular Self-Supervised Learning Framework
Pan Zhou, Yichen Zhou, Chenyang Si +3
In self-supervised learning, multi-granular features are heavily desired though rarely investigated, as different downstream tasks (e.g., general and fine-grained classification) o…
KAN or MLP: A Fairer Comparison
Runpeng Yu, Weihao Yu, Xinchao Wang
This paper does not introduce a novel method. Instead, it offers a fairer and more comprehensive comparison of KAN and MLP models across various tasks, including machine learning,…
Lorentz Equivariant Model for Knowledge-Enhanced Hyperbolic Collaborative Filtering
Bosong Huang, Weihao Yu, Ruzhong Xie +2
Introducing prior auxiliary information from the knowledge graph (KG) to assist the user-item graph can improve the comprehensive performance of the recommender system. Many recent…
MetaFormer Baselines for Vision
Weihao Yu, Chenyang Si, Pan Zhou +5
MetaFormer, the abstracted architecture of Transformer, has been found to play a significant role in achieving competitive performance. In this paper, we further explore the capaci…
LTSP: Long-Term Slice Propagation for Accurate Airway Segmentation
Yangqian Wu, Minghui Zhang, Weihao Yu +3
Purpose: Bronchoscopic intervention is a widely-used clinical technique for pulmonary diseases, which requires an accurate and topological complete airway map for its localization…
ReClor: A Reading Comprehension Dataset Requiring Logical Reasoning
Weihao Yu, Zihang Jiang, Yanfei Dong +1
Recent powerful pre-trained language models have achieved remarkable performance on most of the popular datasets for reading comprehension. It is time to introduce more challenging…
InceptionNeXt: When Inception Meets ConvNeXt
Weihao Yu, Pan Zhou, Shuicheng Yan +1
Inspired by the long-range modeling ability of ViTs, large-kernel convolutions are widely studied and adopted recently to enlarge the receptive field and improve model performance,…
Vision-Language-Action Safety: Threats, Challenges, Evaluations, and Mechanisms
Qi Li, Bo Yin, Weiqi Huang +6
Vision-Language-Action (VLA) models are emerging as a unified substrate for embodied intelligence. This shift raises a new class of safety challenges, stemming from the embodied na…
MonoSplat: Generalizable 3D Gaussian Splatting from Monocular Depth Foundation Models
Yifan Liu, Keyu Fan, Weihao Yu +3
Recent advances in generalizable 3D Gaussian Splatting have demonstrated promising results in real-time high-fidelity rendering without per-scene optimization, yet existing approac…
GeoT: Geometry-guided Instance-dependent Transition Matrix for Semi-supervised Tooth Point Cloud Segmentation
Weihao Yu, Xiaoqing Guo, Chenxin Li +2
Achieving meticulous segmentation of tooth point clouds from intra-oral scans stands as an indispensable prerequisite for various orthodontic applications. Given the labor-intensiv…
SA-3DGS: A Self-Adaptive Compression Method for 3D Gaussian Splatting
Liheng Zhang, Weihao Yu, Zubo Lu +2
Recent advancements in 3D Gaussian Splatting have enhanced efficient and high-quality novel view synthesis. However, representing scenes requires a large number of Gaussian points,…
GTP-4o: Modality-prompted Heterogeneous Graph Learning for Omni-modal Biomedical Representation
Chenxin Li, Xinyu Liu, Cheng Wang +4
Recent advances in learning multi-modal representation have witnessed the success in biomedical domains. While established techniques enable handling multi-modal information, the c…
BREAK: Bronchi Reconstruction by gEodesic transformation And sKeleton embedding
Weihao Yu, Hao Zheng, Minghui Zhang +3
Airway segmentation is critical for virtual bronchoscopy and computer-aided pulmonary disease analysis. In recent years, convolutional neural networks (CNNs) have been widely used…
Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
Li Yuan, Yunpeng Chen, Tao Wang +6
Transformers, which are popular for language modeling, have been explored for solving vision tasks recently, e.g., the Vision Transformer (ViT) for image classification. The ViT mo…
MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities
Weihao Yu, Zhengyuan Yang, Lingfeng Ren +7
MM-Vet, with open-ended vision-language questions targeting at evaluating integrated capabilities, has become one of the most popular benchmarks for large multimodal model evaluati…
Attention Prompting on Image for Large Vision-Language Models
Runpeng Yu, Weihao Yu, Xinchao Wang
Compared with Large Language Models (LLMs), Large Vision-Language Models (LVLMs) can also accept images as input, thus showcasing more interesting emergent capabilities and demonst…
MetaFormer Is Actually What You Need for Vision
Weihao Yu, Mi Luo, Pan Zhou +5
Transformers have shown great potential in computer vision tasks. A common belief is their attention-based token mixer module contributes most to their competence. However, recent…
X-GRM: Large Gaussian Reconstruction Model for Sparse-view X-rays to Computed Tomography
Yifan Liu, Wuyang Li, Weihao Yu +4
Computed Tomography serves as an indispensable tool in clinical workflows, providing non-invasive visualization of internal anatomical structures. Existing CT reconstruction works…
MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Weihao Yu, Zhengyuan Yang, Linjie Li +5
We propose MM-Vet, an evaluation benchmark that examines large multimodal models (LMMs) on complicated multimodal tasks. Recent LMMs have shown various intriguing abilities, such a…
Re-thinking and Re-labeling LIDC-IDRI for Robust Pulmonary Cancer Prediction
Hanxiao Zhang, Xiao Gu, Minghui Zhang +6
The LIDC-IDRI database is the most popular benchmark for lung cancer prediction. However, with subjective assessment from radiologists, nodules in LIDC may have entirely different…
LightFusion: A Light-weighted, Double Fusion Framework for Unified Multimodal Understanding and Generation
Zeyu Wang, Zilong Chen, Chenhui Gou +8
Unified multimodal models have recently shown remarkable gains in both capability and versatility, yet most leading systems are still trained from scratch and require substantial c…
AsyMoE: Leveraging Modal Asymmetry for Enhanced Expert Specialization in Large Vision-Language Models
Heng Zhang, Haichuan Hu, Yaomin Shen +9
Large Vision-Language Models (LVLMs) have demonstrated impressive performance on multimodal tasks through scaled architectures and extensive training. However, existing Mixture of…
Knowledge-Embedded Routing Network for Scene Graph Generation
Tianshui Chen, Weihao Yu, Riquan Chen +1
To understand a scene in depth not only involves locating/recognizing individual objects, but also requires to infer the relationships and interactions among them. However, since t…
ODformer: Spatial-Temporal Transformers for Long Sequence Origin-Destination Matrix Forecasting Against Cross Application Scenario
Jin Huang, Bosong Huang, Weihao Yu +3
Origin-Destination (OD) matrices record directional flow data between pairs of OD regions. The intricate spatiotemporal dependency in the matrices makes the OD matrix forecasting (…
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward
Yuwei Niu, Weiyang Jin, Jiaqi Liao +7
Recent years have witnessed significant progress in Unified Multimodal Models, yet a fundamental question remains: Does understanding truly inform generation? To investigate this,…
Refiner: Refining Self-attention for Vision Transformers
Daquan Zhou, Yujun Shi, Bingyi Kang +6
Vision Transformers (ViTs) have shown competitive accuracy in image classification tasks compared with CNNs. Yet, they generally require much more data for model pre-training. Most…
Deep Reasoning with Knowledge Graph for Social Relationship Understanding
Zhouxia Wang, Tianshui Chen, Jimmy Ren +3
Social relationships (e.g., friends, couple etc.) form the basis of the social network in our daily life. Automatically interpreting such relationships bears a great potential for…
Emerging Properties in Unified Multimodal Pretraining
Chaorui Deng, Deyao Zhu, Kunchang Li +9
Unifying multimodal understanding and generation has shown impressive capabilities in cutting-edge proprietary systems. In this work, we introduce BAGEL, an open-source foundationa…
Heterogeneous Graph Learning for Visual Commonsense Reasoning
Weijiang Yu, Jingwen Zhou, Weihao Yu +2
Visual commonsense reasoning task aims at leading the research field into solving cognition-level reasoning with the ability of predicting correct answers and meanwhile providing c…
LV-BERT: Exploiting Layer Variety for BERT
Weihao Yu, Zihang Jiang, Fei Chen +2
Modern pre-trained language models are mostly built upon backbones stacking self-attention and feed-forward layers in an interleaved order. In this paper, beyond this stereotyped l…
Seed1.5-VL Technical Report
Dong Guo, Faming Wu, Feida Zhu +194
We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…
Inception Transformer
Chenyang Si, Weihao Yu, Pan Zhou +3
Recent studies show that Transformer has strong capability of building long-range dependencies, yet is incompetent in capturing high frequencies that predominantly convey local inf…
ConvBERT: Improving BERT with Span-based Dynamic Convolution
Zihang Jiang, Weihao Yu, Daquan Zhou +3
Pre-trained language models like BERT and its variants have recently achieved impressive performance in various natural language understanding tasks. However, BERT heavily relies o…
X-Gaussian: 4D Radiative Gaussian Splatting for Continuous-time Tomographic Reconstruction
Weihao Yu, Yuanhao Cai, Ruyi Zha +3
Four-dimensional computed tomography (4D CT) reconstruction is crucial for capturing dynamic anatomical changes but faces inherent limitations from conventional phase-binning workf…
MambaOut: Do We Really Need Mamba for Vision?
Weihao Yu, Xinchao Wang
Mamba, an architecture with RNN-like token mixer of state space model (SSM), was recently introduced to address the quadratic complexity of the attention mechanism and subsequently…
NoLan: Mitigating Object Hallucinations in Large Vision-Language Models via Dynamic Suppression of Language Priors
Lingfeng Ren, Weihao Yu, Runpeng Yu +1
Object hallucination is a critical issue in Large Vision-Language Models (LVLMs), where outputs include objects that do not appear in the input image. A natural question arises fro…
Top-Down Compression: Revisit Efficient Vision Token Projection for Visual Instruction Tuning
Bonan li, Zicheng Zhang, Songhua Liu +2
Visual instruction tuning aims to enable large language models to comprehend the visual world, with a pivotal challenge lying in establishing an effective vision-to-language projec…
FDA: Feature Decomposition and Aggregation for Robust Airway Segmentation
Minghui Zhang, Xin Yu, Hanxiao Zhang +5
3D Convolutional Neural Networks (CNNs) have been widely adopted for airway segmentation. The performance of 3D CNNs is greatly influenced by the dataset while the public airway da…
Two-stage Denoising Diffusion Model for Source Localization in Graph Inverse Problems
Bosong Huang, Weihao Yu, Ruzhong Xie +2
Source localization is the inverse problem of graph information dissemination and has broad practical applications. However, the inherent intricacy and uncertainty in information d…
LinFusion: 1 GPU, 1 Minute, 16K Image
Songhua Liu, Weihao Yu, Zhenxiong Tan +1
Modern diffusion models, particularly those utilizing a Transformer-based UNet for denoising, rely heavily on self-attention operations to manage complex spatial relationships, thu…