6 papers
GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios
Yiming Xu, Jihua Kang, Chunsai Du +5
In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient information communication, transformi…
From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes
Qifan Zhang, Sai Haneesh Allu, Jikai Wang +2
Detecting and segmenting novel object instances in open-world environments is a fundamental problem in robotic perception. Given only a small set of template images, a robot must l…
R: A LLM Based Novel-to-Screenplay Generation Framework with Causal Plot Graphs
Zefeng Lin, Yi Xiao, Zhiqiang Mo +8
Automatically adapting novels into screenplays is important for the TV, film, or opera industries to promote products with low costs. The strong performances of large language mode…
Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting
Qifan Zhang, Yunhui Guo, Yu Xiang
Prompt-based continual learning has shown strong performance in rehearsal-free class-incremental learning by adapting learnable prompts while freezing a pre-trained Vision Transfor…
HO-Cap: A Capture System and Dataset for 3D Reconstruction and Pose Tracking of Hand-Object Interaction
Jikai Wang, Qifan Zhang, Yu-Wei Chao +3
We introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos. The system leverages multiple RGBD cameras and…
CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities
Rohith Peddi, Shivvrat Arya, Bharath Challa +10
Following step-by-step procedures is an essential component of various activities carried out by individuals in their daily lives. These procedures serve as a guiding framework tha…