12 papers
Musical Score Understanding Benchmark: Evaluating Large Language Models' Comprehension of Complete Musical Scores
Congren Dai, Yue Yang, Krinos Li +12
Understanding complete musical scores entails integrated reasoning over pitch, rhythm, harmony, and large-scale structure, yet the ability of Large Language Models and Vision--Lang…
How Adversarial Environments Mislead Agentic AI?
Zhonghao Zhan, Huichi Zhou, Zhenhao Li +3
Tool-integrated agents are deployed on the premise that external tools ground their outputs in reality. Yet this very reliance creates a critical attack surface. Current evaluation…
Seeing Through Experts Eyes A Foundational Vision Language Model Trained on Radiologists Gaze and Reasoning
Kinhei Lee, Peiyuan Jing, Zhenxuan Zhang +5
Large scale vision language models have shown promise in automating chest Xray interpretation, yet their clinical utility remains limited by a gap between model outputs and radiolo…
CT-Conditioned Diffusion Prior with Physics-Constrained Sampling for PET Super-Resolution
Liutao Yang, Zi Wang, Peiyuan Jing +5
PET super-resolution is highly under-constrained because paired multi-resolution scans from the same subject are rarely available, and effective resolution is determined by scanner…
ReDiff: Reliability-Guided Diffusion for Trustworthy Ultra-Low-Field to High-Field MRI Synthesis
Zhenxuan Zhang, Peiyuan Jing, Ruicheng Yuan +9
Low-field to high-field MRI synthesis has emerged as a promising strategy to improve image quality when access to high-field scanners is limited. However, in ultra-low-field settin…
ProSMA-UNet: Decoder Conditioning for Proximal-Sparse Skip Feature Selection
Chun-Wun Cheng, Yanqi Cheng, Peiyuan Jing +4
Medical image segmentation commonly relies on U-shaped encoder-decoder architectures such as U-Net, where skip connections preserve fine spatial detail by injecting high-resolution…