4 papers
From Pixels to Words -- Towards Native One-Vision Models at Scale
Haiwen Diao, Jiahao Wang, Penghao Wu +18
Current vision-language models (VLMs) typically stitch together separate image encoders and language decoders via multi-stage alignment, a modular framework that inevitably fragmen…
From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
Haiwen Diao, Mingxuan Li, Silei Wu +6
The edifice of native Vision-Language Models (VLMs) has emerged as a rising contender to typical modular VLMs, shaped by evolving model architectures and training paradigms. Yet, t…
IFG: Internet-Scale Guidance for Functional Grasping Generation
Ray Muxin Liu, Mingxuan Li, Kenneth Shaw +1
Large Vision Models trained on internet-scale data have demonstrated strong capabilities in segmenting and semantically understanding object parts, even in cluttered, crowded scene…
HDCNet: A Hybrid Depth Completion Network for Grasping Transparent and Reflective Objects
Guanghu Xie, Mingxu Li, Songwei Wu +4
Depth perception of transparent and reflective objects has long been a critical challenge in robotic manipulation.Conventional depth sensors often fail to provide reliable measurem…