1 citations · 2 across the 31 of their papers we have counts for
6 papers · 1 filter
ROOT: VLM based System for Indoor Scene Understanding and Beyond
Yonghui Wang, Shi-Yong Chen, Zhenxing Zhou +4
Recently, Vision Language Models (VLMs) have experienced significant advancements, yet these models still face challenges in spatial hierarchical reasoning within indoor scenes. In…
Make-It-Animatable: An Efficient Framework for Authoring Animation-Ready 3D Characters
Zhiyang Guo, Jinxu Xiang, Kai Ma +3
3D characters are essential to modern creative industries, but making them animatable often demands extensive manual work in tasks like rigging and skinning. Existing automatic rig…
LaneTCA: Enhancing Video Lane Detection with Temporal Context Aggregation
Keyi Zhou, Li Li, Wengang Zhou +3
In video lane detection, there are rich temporal contexts among successive frames, which is under-explored in existing lane detectors. In this work, we propose LaneTCA to bridge th…
SwinShadow: Shifted Window for Ambiguous Adjacent Shadow Detection
Yonghui Wang, Shaokai Liu, Li Li +2
Shadow detection is a fundamental and challenging task in many computer vision applications. Intuitively, most shadows come from the occlusion of light by the object itself, result…
Forest2Seq: Revitalizing Order Prior for Sequential Indoor Scene Synthesis
Qi Sun, Hang Zhou, Wengang Zhou +2
Synthesizing realistic 3D indoor scenes is a challenging task that traditionally relies on manual arrangement and annotation by expert designers. Recent advances in autoregressive…
TabPedia: Towards Comprehensive Visual Table Understanding with Concept Synergy
Weichao Zhao, Hao Feng, Qi Liu +9
Tables contain factual and quantitative data accompanied by various structures and contents that pose challenges for machine comprehension. Previous methods generally design task-s…