4 papers
HyperET: Efficient Training in Hyperbolic Space for Multi-modal Large Language Models
Zelin Peng, Zhengqin Xu, Qingyang Liu +2
Multi-modal large language models (MLLMs) have emerged as a transformative approach for aligning visual and textual understanding. They typically require extremely high computation…
Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
Yu Huang, Zelin Peng, Changsong Wen +2
Affordance segmentation aims to decompose 3D objects into parts that serve distinct functional roles, enabling models to reason about object interactions rather than mere recogniti…
AnimateScene: Camera-controllable Animation in Any Scene
Qingyang Liu, Bingjie Gao, Weiheng Huang +10
Recent advances in 3D scene reconstruction and 4D human animation have broadened adoption, but integrating the two remains difficult. Key challenges include placing humans at plaus…
MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
Yu Huang, Zelin Peng, Yichen Zhao +3
Medical image segmentation is crucial for clinical diagnosis, yet existing models are limited by their reliance on explicit human instructions and lack the active reasoning capabil…