9 papers
Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation
Shun Liu, Nan Xi, Yang Liu +3
Referring Image Segmentation (RIS) aims to segment image regions specified by natural language, enabling fine-grained and controllable visual understanding. Extending RIS to endosc…
FairLLaVA: Fairness-Aware Parameter-Efficient Fine-Tuning for Large Vision-Language Assistants
Mahesh Bhosale, Abdul Wasi, Shantam Srivastava +5
While powerful in image-conditioned generation, multimodal large language models (MLLMs) can display uneven performance across demographic groups, highlighting fairness risks. In s…
PP-Motion: Physical-Perceptual Fidelity Evaluation for Human Motion Generation
Sihan Zhao, Zixuan Wang, Tianyu Luan +5
Human motion generation has found widespread applications in AR/VR, film, sports, and medical rehabilitation, offering a cost-effective alternative to traditional motion capture sy…
VLM-Guided Iterative Refinement for Surgical Image Segmentation with Foundation Models
Ange Lou, Yamin Li, Qi Chang +4
Surgical image segmentation is essential for robot-assisted surgery and intraoperative guidance. However, existing methods are constrained to predefined categories, produce one-sho…
Textured Geometry Evaluation: Perceptual 3D Textured Shape Metric via 3D Latent-Geometry Network
Tianyu Luan, Xuelu Feng, Zixin Zhu +6
Textured high-fidelity 3D models are crucial for games, AR/VR, and film, but human-aligned evaluation methods still fall behind despite recent advances in 3D reconstruction and gen…
SRAM: Shape-Realism Alignment Metric for No Reference 3D Shape Evaluation
Sheng Liu, Tianyu Luan, Phani Nuney +2
3D generation and reconstruction techniques have been widely used in computer games, film, and other content creation areas. As the application grows, there is a growing demand for…