papers

Publications (89)

physics.app-ph2019

Intelligent Metasurface Imager and Recognizer

Lianlin Li, Ya Shuang, Qian Ma +7

It is ever-increasingly demanded to remotely monitor people in daily life using radio-frequency probing signals. However, conventional systems can hardly be deployed in real-world…

cs.CL2024

EVOR: Evolving Retrieval for Code Generation

Hongjin Su, Shuyang Jiang, Yuhang Lai +5

Recently the retrieval-augmented generation (RAG) has been successfully applied in code generation. However, existing pipelines for retrieval-augmented code generation (RACG) emplo…

cs.CV2024

G2D: From Global to Dense Radiography Representation Learning via Vision-Language Pre-training

Che Liu, Cheng Ouyang, Sibo Cheng +3

Recently, medical vision-language pre-training (VLP) has reached substantial progress to learn global visual representation from medical images and their paired radiology reports.…

cs.CL2024

Parrot: Enhancing Multi-Turn Instruction Following for Large Language Models

Yuchong Sun, Che Liu, Kun Zhou +6

Humans often interact with large language models (LLMs) in multi-turn interaction to obtain desired answers or more information. However, most existing studies overlook the multi-t…

cs.LG2023

Frozen Language Model Helps ECG Zero-Shot Learning

Jun Li, Che Liu, Sibo Cheng +2

The electrocardiogram (ECG) is one of the most commonly used non-invasive, convenient medical monitoring tools that assist in the clinical diagnosis of heart diseases. Recently, de…

cs.CL2022

Dial2vec: Self-Guided Contrastive Learning of Unsupervised Dialogue Embeddings

Che Liu, Rui Wang, Junfeng Jiang +2

In this paper, we introduce the task of learning unsupervised dialogue embeddings. Trivial approaches such as combining pre-trained word or sentence embeddings and encoding through…

eess.IV2025

Towards Cardiac MRI Foundation Models: Comprehensive Visual-Tabular Representations for Whole-Heart Assessment and Beyond

Yundi Zhang, Paul Hager, Che Liu +4

Cardiac magnetic resonance imaging is the gold standard for non-invasive cardiac assessment, offering rich spatio-temporal views of the cardiac anatomy and physiology. Patient-leve…

cs.AI2026

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +306

Recent AI systems have achieved strong results on a wide range of benchmarks, yet these gains have not translated into economically meaningful deployment across many professional d…

cs.SD2024

Reverse the auditory processing pathway: Coarse-to-fine audio reconstruction from fMRI

Che Liu, Changde Du, Xiaoyu Chen +1

Drawing inspiration from the hierarchical processing of the human auditory system, which transforms sound from low-level acoustic features to high-level semantic understanding, we…

cs.CL2023

OpenAgents: An Open Platform for Language Agents in the Wild

Tianbao Xie, Fan Zhou, Zhoujun Cheng +13

Language agents show potential in being capable of utilizing natural language for varied and intricate tasks in diverse environments, particularly when built upon large language mo…

cs.CV2025

Dual form Complementary Masking for Domain-Adaptive Image Segmentation

Jiawen Wang, Yinda Chen, Xiaoyu Liu +4

Recent works have correlated Masked Image Modeling (MIM) with consistency regularization in Unsupervised Domain Adaptation (UDA). However, they merely treat masking as a special fo…

cs.CL2025

Beyond Distillation: Pushing the Limits of Medical LLM Reasoning with Minimalist Rule-Based RL

Che Liu, Haozhe Wang, Jiazhen Pan +6

Improving performance on complex tasks and enabling interpretable decision making in large language models (LLMs), especially for clinical applications, requires effective reasonin…

cs.AI2025

Emergent Hierarchical Reasoning in LLMs through Reinforcement Learning

Haozhe Wang, Qixin Xu, Che Liu +3

Reinforcement Learning (RL) has proven highly effective at enhancing the complex reasoning abilities of Large Language Models (LLMs), yet underlying mechanisms driving this success…

eess.SP2025

SuPreME: A Supervised Pre-training Framework for Multimodal ECG Representation Learning

Mingsheng Cai, Jiuming Jiang, Wenhao Huang +2

Cardiovascular diseases are a leading cause of death and disability worldwide. Electrocardiogram (ECG) is critical for diagnosing and monitoring cardiac health, but obtaining large…

cs.CV2025

Can Medical Vision-Language Pre-training Succeed with Purely Synthetic Data?

Che Liu, Zhongwei Wan, Haozhe Wang +6

Medical Vision-Language Pre-training (MedVLP) has made significant progress in enabling zero-shot tasks for medical image understanding. However, training MedVLP models typically r…

cs.CV2024

Utilizing Synthetic Data for Medical Vision-Language Pre-training: Bypassing the Need for Real Images

Che Liu, Anand Shah, Wenjia Bai +1

Medical Vision-Language Pre-training (VLP) learns representations jointly from medical images and paired radiology reports. It typically requires large-scale paired image-text data…

cs.LG2025

Pelican-VL 1.0: A Foundation Brain Model for Embodied Intelligence

Yi Zhang, Che Liu, Xiancong Ren +20

This report presents Pelican-VL 1.0, a new family of open-source embodied brain models with parameter scales ranging from 7 billion to 72 billion. Our explicit mission is clearly s…

cs.CL2025

Step-Audio 2 Technical Report

Boyong Wu, Chao Yan, Chen Hu +106

This paper presents Step-Audio 2, an end-to-end multi-modal large language model designed for industry-strength audio understanding and speech conversation. By integrating a latent…

cs.CV2025

BOTM: Echocardiography Segmentation via Bi-directional Optimal Token Matching

Zhihua Liu, Lei Tong, Xilin He +4

Existed echocardiography segmentation methods often suffer from anatomical inconsistency challenge caused by shape variation, partial observation and region ambiguity with similar…

cs.CV2026

MedOpenClaw and MedFlowBench: Auditing Medical Agents in Full-Study Workflows

Weixiang Shen, Chengzhi Shen, Yanzhu Hu +12

Medical imaging benchmarks often evaluate VLMs on pre-selected 2D images, slices, crops, or patches, making evaluation closer to visual recognition. Real clinical workflows impose…

cs.LG2026

GDSD: Reinforcement Learning as Guided Denoiser Self-Distillation for Diffusion Language Models

Xiaohang Tang, Keyue Jiang, Che Liu +4

Reinforcement learning (RL) can be used to improve the policy (denoiser) of diffusion large language models (dLLMs), while being hindered by the intractability of the policy likeli…

cs.LG2025

Knowledge-enhanced Multimodal ECG Representation Learning with Arbitrary-Lead Inputs

Che Liu, Cheng Ouyang, Zhongwei Wan +3

Recent advances in multimodal ECG representation learning center on aligning ECG signals with paired free-text reports. However, suboptimal alignment persists due to the complexity…

quant-ph2025

Spin squeezing in an ensemble of nitrogen-vacancy centers in diamond

Weijie Wu, Emily J. Davis, Lillian B. Hughes +10

Spin squeezed states provide a seminal example of how the structure of quantum mechanical correlations can be controlled to produce metrologically useful entanglement. Such squeeze…

cs.CV2024

FMBench: Benchmarking Fairness in Multimodal Large Language Models on Medical Tasks

Peiran Wu, Che Liu, Canyu Chen +3

Advancements in Multimodal Large Language Models (MLLMs) have significantly improved medical task performance, such as Visual Question Answering (VQA) and Report Generation (RG). H…

cs.MM2026

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

Che Liu, Lichao Ma, Xiangyu Tony Zhang +4

Omni-modal language models are intended to jointly understand audio, visual inputs, and language, but benchmark gains can be inflated when visual evidence alone is enough to answer…

cs.MM2025

Nexus: An Omni-Perceptive And -Interactive Model for Language, Audio, And Vision

Che Liu, Yingji Zhang, Dong Zhang +13

This work proposes an industry-level omni-modal large language model (LLM) pipeline that integrates auditory, visual, and linguistic modalities to overcome challenges such as limit…

cs.CV2025

Argus: Benchmarking and Enhancing Vision-Language Models for 3D Radiology Report Generation

Che Liu, Zhongwei Wan, Yuqi Wang +5

Automatic radiology report generation holds significant potential to streamline the labor-intensive process of report writing by radiologists, particularly for 3D radiographs such…

eess.SP2024

Zero-Shot ECG Classification with Multimodal Learning and Test-time Clinical Knowledge Enhancement

Che Liu, Zhongwei Wan, Cheng Ouyang +3

Electrocardiograms (ECGs) are non-invasive diagnostic tools crucial for detecting cardiac arrhythmic diseases in clinical practice. While ECG Self-supervised Learning (eSSL) method…

cs.CV2023

M-FLAG: Medical Vision-Language Pre-training with Frozen Language Models and Latent Space Geometry Optimization

Che Liu, Sibo Cheng, Chen Chen +5

Medical vision-language models enable co-learning and integrating features from medical imaging and clinical text. However, these models are not easy to train and the latent repres…

eess.IV2025

NOVA: A Benchmark for Anomaly Localization and Clinical Reasoning in Brain MRI

Cosmin I. Bercea, Jun Li, Philipp Raffler +12

In many real-world applications, deployed models encounter inputs that differ from the data seen during training. Out-of-distribution detection identifies whether an input stems fr…

cs.AI2025

Bridging VLMs and Embodied Intelligence with Deliberate Practice Policy Optimization

Yi Zhang, Che Liu, Xiancong Ren +17

Developing a universal and versatile embodied intelligence system presents two primary challenges: the critical embodied data bottleneck, where real-world data is scarce and expens…

cs.CV2025

Knowledge to Sight: Reasoning over Visual Attributes via Knowledge Decomposition for Abnormality Grounding

Jun Li, Che Liu, Wenjia Bai +4

In this work, we address the problem of grounding abnormalities in medical images, where the goal is to localize clinical findings based on textual descriptions. While generalist V…

cs.CV2025

M3Ret: Unleashing Zero-shot Multimodal Medical Image Retrieval via Self-Supervision

Che Liu, Zheng Jiang, Chengyu Fang +5

Medical image retrieval is essential for clinical decision-making and translational research, relying on discriminative visual representations. Yet, current methods remain fragment…

cs.CL2025

SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning

Zhongwei Wan, Zhihao Dou, Che Liu +11

Multimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle with complex problems requiring explicit self-reflection and self-…

cs.CV2025

PlanMoGPT: Flow-Enhanced Progressive Planning for Text to Motion Synthesis

Chuhao Jin, Haosen Li, Bingzi Zhang +7

Recent advances in large language models (LLMs) have enabled breakthroughs in many multimodal generation tasks, but a significant performance gap still exists in text-to-motion gen…

cs.CL2024

Efficient Large Language Models: A Survey

Zhongwei Wan, Xin Wang, Che Liu +9

Large Language Models (LLMs) have demonstrated remarkable capabilities in important tasks such as natural language understanding and language generation, and thus have the potentia…

cs.RO2026

Artificial Foveated Perception for Mitigating Shortcut Learning in Robotic Foundation Models

Xiatao Sun, Yuan Zhuang, Mateo Sanchez Lopez Negrete +9

Robotic foundation models have recently made substantial progress in multi-task capability, cross-embodiment transfer, and language-conditioned control. Yet robust deployment acros…

cs.LG2025

Machine learning for modelling unstructured grid data in computational physics: a review

Sibo Cheng, Marc Bocquet, Weiping Ding +20

Unstructured grid data are essential for modelling complex geometries and dynamics in computational physics. Yet, their inherent irregularity presents significant challenges for co…

cs.CL2024

Enhancing Role-playing Systems through Aggressive Queries: Evaluation and Improvement

Yihong Tang, Jiao Ou, Che Liu +3

The advent of Large Language Models (LLMs) has propelled dialogue generation into new realms, particularly in the field of role-playing systems (RPSs). While enhanced with ordinary…

cs.IR2018

DeepNIS: Deep Neural Network for Nonlinear Electromagnetic Inverse Scattering

Lianlin Li, Long Gang Wang, Fernando L. Teixeira +3

Nonlinear electromagnetic (EM) inverse scattering is a quantitative and super-resolution imaging technique, in which more realistic interactions between the internal structure of s…

cs.CV2024

BIMCV-R: A Landmark Dataset for 3D CT Text-Image Retrieval

Yinda Chen, Che Liu, Xiaoyu Liu +2

The burgeoning integration of 3D medical imaging into healthcare has led to a substantial increase in the workload of medical professionals. To assist clinicians in their diagnosti…

cs.AI2026

Evo-MedAgent: Beyond One-Shot Diagnosis with Agents That Remember, Reflect, and Improve

Weixiang Shen, Bailiang Jian, Jun Li +6

Tool-augmented large language model (LLM) agents can orchestrate specialist classifiers, segmentation models, and visual question-answering modules to interpret chest X-rays. Howev…

cs.CV2026

Does DINOv3 Set a New Medical Vision Standard? Benchmarking 2D and 3D Classification, Segmentation, and Registration

Che Liu, Yinda Chen, Haoyuan Shi +21

The advent of large-scale vision foundation models, pre-trained on diverse natural images, has marked a paradigm shift in computer vision. However, how the frontier vision foundati…

cs.CV2024

MME-Finance: A Multimodal Finance Benchmark for Expert-level Understanding and Reasoning

Ziliang Gan, Yu Lu, Dong Zhang +9

In recent years, multimodal benchmarks for general domains have guided the rapid development of multimodal models on general tasks. However, the financial field has its peculiariti…

cs.CE2024

PointEMRay: A Novel Efficient SBR Framework on Point Based Geometry

Kaiqiao Yang, Che Liu, Wenming Yu +1

The rapid computation of electromagnetic (EM) fields across various scenarios has long been a challenge, primarily due to the need for precise geometric models. The emergence of po…

cs.CL2024

LOOK-M: Look-Once Optimization in KV Cache for Efficient Multimodal Long-Context Inference

Zhongwei Wan, Ziang Wu, Che Liu +5

Long-context Multimodal Large Language Models (MLLMs) demand substantial computational resources for inference as the growth of their multimodal Key-Value (KV) cache, in response t…

cs.LG2023

Machine learning with data assimilation and uncertainty quantification for dynamical systems: a review

Sibo Cheng, Cesar Quilodran-Casas, Said Ouala +14

Data Assimilation (DA) and Uncertainty quantification (UQ) are extensively used in analysing and reducing error propagation in high-dimensional spatial-temporal dynamics. Typical a…

cs.CL2025

MEIT: Multimodal Electrocardiogram Instruction Tuning on Large Language Models for Report Generation

Zhongwei Wan, Che Liu, Xin Wang +6

Electrocardiogram (ECG) is the primary non-invasive diagnostic tool for monitoring cardiac conditions and is crucial in assisting clinicians. Recent studies have concentrated on cl…

cs.CL2026

Dynamic Decision Learning: Test-Time Evolution for Abnormality Grounding in Rare Diseases

Jun Li, Mingxuan Liu, Jiazhen Pan +4

Clinical abnormality grounding for rare diseases is often hindered by data scarcity, making supervised fine-tuning impractical and single-pass inference highly unstable. We propose…

cs.CL2024

ERABAL: Enhancing Role-Playing Agents through Boundary-Aware Learning

Yihong Tang, Jiao Ou, Che Liu +3

Role-playing is an emerging application in the field of Human-Computer Interaction (HCI), primarily implemented through the alignment training of a large language model (LLM) with…

cs.IT2022

Directly wireless communication of human minds via non-invasive brain-computer-metasurface platform

Qian Ma, Wei Gao, Qiang Xiao +13

Brain-computer interfaces (BCIs), invasive or non-invasive, have projected unparalleled vision and promise for assisting patients in need to better their interaction with the surro…

cs.CL2024

Inductive-Deductive Strategy Reuse for Multi-Turn Instructional Dialogues

Jiao Ou, Jiayu Wu, Che Liu +3

Aligning large language models (LLMs) with human expectations requires high-quality instructional dialogues, which usually require instructions that are diverse and in-depth. Exist…

cs.CV2023

Generative Text-Guided 3D Vision-Language Pretraining for Unified Medical Image Segmentation

Yinda Chen, Che Liu, Wei Huang +3

Vision-Language Pretraining (VLP) has demonstrated remarkable capabilities in learning visual representations from textual descriptions of images without annotations. Yet, effectiv…

cond-mat.soft2026

Programming strain-stiffening in soft composites via structural memory near jamming

Yiqiu Zhao, Deng Pan, Yiming Pang +6

Soft composite solids, comprising discrete inclusions embedded within a compliant matrix, are emerging candidates for engineering synthetic tissues and soft robotic materials. Curr…

eess.SP2026

Metasurface embodied intelligence through electromagnetic world model

Che Liu, Zhenhao Fu, Qian Ma +4

Mastering invisible electromagnetic (EM) environment and sculpting radio waves with the dexterity of manipulating light or matter have long been aspirations in physics and informat…

cs.CV2026

Search Beyond What Can Be Taught: Evolving the Knowledge Boundary in Agentic Visual Generation

Haozhe Wang, Weijia Feng, Jinpeng Yu +8

Visual generators excel at rendering, but they confidently fabricate what they do not know. User requests are unbounded, evolving, and deeply long-tailed: new characters, trending…

cs.CV2024

Freeze the backbones: A Parameter-Efficient Contrastive Approach to Robust Medical Vision-Language Pre-training

Jiuming Qin, Che Liu, Sibo Cheng +2

Modern healthcare often utilises radiographic images alongside textual reports for diagnostics, encouraging the use of Vision-Language Self-Supervised Learning (VL-SSL) with large…

cs.LG2025

An Electrocardiogram Foundation Model Built on over 10 Million Recordings with External Evaluation across Multiple Domains

Jun Li, Aaron Aguirre, Junior Moura +6

Artificial intelligence (AI) has demonstrated significant potential in ECG analysis and cardiovascular disease assessment. Recently, foundation models have played a remarkable role…

cs.MS2024

TorchDA: A Python package for performing data assimilation with deep learning forward and transformation functions

Sibo Cheng, Jinyang Min, Che Liu +1

Data assimilation techniques are often confronted with challenges handling complex high dimensional physical systems, because high precision simulation in complex high dimensional…

cs.CV2025

MedVLM-R1: Incentivizing Medical Reasoning Capability of Vision-Language Models (VLMs) via Reinforcement Learning

Jiazhen Pan, Che Liu, Junde Wu +6

Reasoning is a critical frontier for advancing medical image analysis, where transparency and trustworthiness play a central role in both clinician trust and regulatory approval. A…

cs.CL2024

DialogBench: Evaluating LLMs as Human-like Dialogue Systems

Jiao Ou, Junda Lu, Che Liu +4

Large language models (LLMs) have achieved remarkable breakthroughs in new dialogue capabilities by leveraging instruction tuning, which refreshes human impressions of dialogue sys…

eess.SP2023

ETP: Learning Transferable ECG Representations via ECG-Text Pre-training

Che Liu, Zhongwei Wan, Sibo Cheng +2

In the domain of cardiovascular healthcare, the Electrocardiogram (ECG) serves as a critical, non-invasive diagnostic tool. Although recent strides in self-supervised learning (SSL…

cs.LG2026

PPGFlowECG: Latent Rectified Flow with Cross-Modal Encoding for PPG-Guided ECG Generation and Cardiovascular Disease Detection

Xiaocheng Fang, Jiarui Jin, Haoyu Wang +8

Electrocardiography (ECG) is the clinical gold standard for cardiovascular disease (CVD) assessment, yet continuous monitoring is constrained by the need for dedicated hardware and…

cs.HC2025

BP-GPT: Auditory Neural Decoding Using fMRI-prompted LLM

Xiaoyu Chen, Changde Du, Che Liu +2

Decoding language information from brain signals represents a vital research area within brain-computer interfaces, particularly in the context of deciphering the semantic informat…

cs.CV2026

Unified Multimodal Model for Brain MRI Imputation and Understanding

Zhiyun Song, Che Liu, Tian Xia +2

Multimodal large language models (MLLMs) hold great potential for medicine, as they inherit knowledge from LLM and allow multiple data modalities to be integrated, analysed and int…

cs.CL2021

DialogueCSE: Dialogue-based Contrastive Learning of Sentence Embeddings

Che Liu, Rui Wang, Jinghua Liu +3

Learning sentence embeddings from dialogues has drawn increasing attention due to its low annotation cost and high domain adaptability. Conventional approaches employ the siamese-n…

cs.AI2026

Does RLVR Extend Reasoning Boundaries? Investigating Capability Expansion in Vision-Language Models

Minghe Shen, Zhuo Zhi, Chonghan Liu +3

Recent studies posit that Reinforcement Learning with Verifiable Rewards (RLVR) primarily amplifies behaviors inherent to the pre-training distribution rather than inducing new cap…

cs.CL2025

MEDA: Dynamic KV Cache Allocation for Efficient Multimodal Long-Context Inference

Zhongwei Wan, Hui Shen, Xin Wang +3

Long-context Multimodal Large Language Models (MLLMs) that incorporate long text-image and text-video modalities, demand substantial resources as their multimodal Key-Value (KV) ca…

cs.LG2023

Spectral Cross-Domain Neural Network with Soft-adaptive Threshold Spectral Enhancement

Che Liu, Sibo Cheng, Weiping Ding +1

Electrocardiography (ECG) signals can be considered as multi-variable time-series. The state-of-the-art ECG data classification approaches, based on either feature engineering or d…

cs.RO2026

Pelican-Unify 1.0: A Unified Embodied Intelligence Model for Understanding, Reasoning, Imagination and Action

Yi Zhang, Yinda Chen, Che Liu +26

We present Pelican-Unify 1.0, the first embodied foundation model trained according to the principle of unification. Pelican-Unify 1.0 uses a single VLM as a unified understanding…

cs.LG2026

Scaling World-Model Reinforcement Learning Through Diffusion Policy Optimization

Xiaoyuan Cheng, Wenxuan Yuan, Zhancun Mu +5

Model-based reinforcement learning (RL) can be effectively supported at scale through the use of world models. However, in practice, scaling such approaches remains fundamentally l…

eess.SP2025

Fine-grained Contrastive Learning for ECG-Report Alignment with Waveform Enhancement

Haitao Li, Che Liu, Zhengyao Ding +3

Electrocardiograms (ECGs) are essential for diagnosing cardiovascular diseases. However, existing ECG-Report contrastive learning methods focus on whole-ECG and report alignment, m…

cs.CV2025

CogDoc: Towards Unified thinking in Documents

Qixin Xu, Haozhe Wang, Che Liu +2

Current document reasoning paradigms are constrained by a fundamental trade-off between scalability (processing long-context documents) and fidelity (capturing fine-grained, multim…

eess.AS2026

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action

Haoyang Zhang, Jun Chen, Donghang Wu +13

Recent advances in spoken dialogue language models have shifted from turn-based to full-duplex designs, where the model continuously listens to the user while generating responses.…

cs.CV2024

How Does Diverse Interpretability of Textual Prompts Impact Medical Vision-Language Zero-Shot Tasks?

Sicheng Wang, Che Liu, Rossella Arcucci

Recent advancements in medical vision-language pre-training (MedVLP) have significantly enhanced zero-shot medical vision tasks such as image classification by leveraging large-sca…

cs.AI2026

IQuest-Coder-V1 Technical Report

Jian Yang, Wei Zhang, Shawn Guo +35

In this report, we introduce the IQuest-Coder-V1 series-(7B/14B/40B/40B-Loop), a new family of code large language models (LLMs). Moving beyond static code representations, we prop…

cs.HC2024

Open-vocabulary Auditory Neural Decoding Using fMRI-prompted LLM

Xiaoyu Chen, Changde Du, Che Liu +2

Decoding language information from brain signals represents a vital research area within brain-computer interfaces, particularly in the context of deciphering the semantic informat…

cs.CL2020

Sequential Sentence Matching Network for Multi-turn Response Selection in Retrieval-based Chatbots

Chao Xiong, Che Liu, Zijun Xu +2

Recently, open domain multi-turn chatbots have attracted much interest from lots of researchers in both academia and industry. The dominant retrieval-based methods use context-resp…

cs.CV2025

Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions

Jun Li, Che Liu, Wenjia Bai +3

Visual Language Models (VLMs) have demonstrated impressive capabilities in visual grounding tasks. However, their effectiveness in the medical domain, particularly for abnormality…

cs.CL2026

ClinicalBench: Can LLMs Beat Traditional ML Models in Clinical Prediction?

Canyu Chen, Jian Yu, Shan Chen +8

Large Language Models (LLMs) hold great promise to revolutionize current clinical systems for their superior capacities on medical text processing tasks and medical licensing exams…

math.NT2025

On Weak Approximation of Reductive Groups over Higher Dimensional Function Fields

Zhongda Li, Che Liu, Haoxiang Pan

Let be a -local field of characteristic 0, and let be the function field of a nice curve over . We give a defect to weak approximation for reductive groups over u…

cs.CV2025

How Far Have Medical Vision-Language Models Come? A Comprehensive Benchmarking Study

Che Liu, Jiazhen Pan, Weixiang Shen +3

Vision-Language Models (VLMs) trained on web-scale corpora excel at natural image tasks and are increasingly repurposed for healthcare; however, their competence in medical tasks r…

cs.CV2025

T3D: Advancing 3D Medical Vision-Language Pre-training by Learning Multi-View Visual Consistency

Che Liu, Cheng Ouyang, Yinda Chen +7

While 3D visual self-supervised learning (vSSL) shows promising results in capturing visual representations, it overlooks the clinical knowledge from radiology reports. Meanwhile,…

cs.CL2024

Med-UniC: Unifying Cross-Lingual Medical Vision-Language Pre-Training by Diminishing Bias

Zhongwei Wan, Che Liu, Mi Zhang +6

The scarcity of data presents a critical obstacle to the efficacy of medical visionlanguage pre-training (VLP). A potential solution lies in the combination of datasets from variou…

eess.SP2025

Self-Alignment Learning to Improve Myocardial Infarction Detection from Single-Lead ECG

Jiarui Jin, Xiaocheng Fang, Haoyu Wang +5

Myocardial infarction is a critical manifestation of coronary artery disease, yet detecting it from single-lead electrocardiogram (ECG) remains challenging due to limited spatial i…

cs.LG2023

Efficient deep data assimilation with sparse observations and time-varying sensors

Sibo Cheng, Che Liu, Yike Guo +1

Variational Data Assimilation (DA) has been broadly used in engineering problems for field reconstruction and prediction by performing a weighted combination of multiple sources of…

cs.CV2023

Image Recognition of Oil Leakage Area Based on Logical Semantic Discrimination

Weiying Lin, Che Liu, Xin Zhang +3

Implementing precise detection of oil leaks in peak load equipment through image analysis can significantly enhance inspection quality and ensure the system's safety and reliabilit…

cs.LG2026

Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Jiazhen Pan, Bailiang Jian, Paul Hager +19

The paper presents a dynamic red‑teaming framework (DAS) that continuously stress‑tests large language models on health tasks for robustness, privacy, bias, and hallucination, reve…

#large language models#health AI#safety evaluation#red teaming
cs.CV2024

IMITATE: Clinical Prior Guided Hierarchical Vision-Language Pre-training

Che Liu, Sibo Cheng, Miaojing Shi +3

In the field of medical Vision-Language Pre-training (VLP), significant efforts have been devoted to deriving text and image features from both clinical reports and associated medi…