Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
RSAgent: Learning to Reason and Act for Text-Guided Segmentation via Multi-Turn Tool Invocations
Xingqi He, Yujie Zhang, Shuyong Gao +6
Text-guided object segmentation requires both cross-modal reasoning and pixel grounding abilities. Most recent methods treat text-guided segmentation as one-shot grounding, where t…
cs.CV2025
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark
Ziming Cheng, Binrui Xu, Lisheng Gong +14
With enhanced capabilities and widespread applications, Multimodal Large Language Models (MLLMs) are increasingly required to process and reason over multiple images simultaneously…
cs.CV2024
3D Shape Augmentation with Content-Aware Shape Resizing
Mingxiang Chen, Jian Zhang, Boli Zhou +1
Recent advancements in deep learning for 3D models have propelled breakthroughs in generation, detection, and scene understanding. However, the effectiveness of these algorithms hi…