8 papers
NEARL: Interacted Query Adaptation with Orthogonal Regularization for Medical Vision-Language Understanding
Zelin Peng, Yichen Zhao, Yu Huang +5
Computer-aided medical image analysis is crucial for disease diagnosis and treatment planning. While vision-language models (VLMs) such as CLIP exhibit strong generalization abilit…
Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression
Shuochen Chang, Qingyang Liu, Shaobo Wang +8
Large language models achieve high reasoning performance via explicit chain-of-thought and reinforcement learning, but require long output sequences and extended inference time. La…
MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation
Yichen Zhao, Zelin Peng, Fenghe Tang +3
Chest X-ray (CXR) reporting follows a region-based clinical workflow in which radiologists inspect anatomical regions and integrate localized findings into a final report. However,…
Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
Yu Huang, Zelin Peng, Changsong Wen +2
Affordance segmentation aims to decompose 3D objects into parts that serve distinct functional roles, enabling models to reason about object interactions rather than mere recogniti…
AnimateScene: Camera-controllable Animation in Any Scene
Qingyang Liu, Bingjie Gao, Weiheng Huang +10
Recent advances in 3D scene reconstruction and 4D human animation have broadened adoption, but integrating the two remains difficult. Key challenges include placing humans at plaus…
HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
Zelin Peng, Zhengqin Xu, Qingyang Liu +2
Multi-modal large language models (MLLMs) have emerged as a transformative approach for aligning visual and textual understanding. They typically require extremely high computation…