activity
20242026
collaborators

6 papers

cs.LG2026

GVR-Coder: A Visual-Feedback Framework for Structured SVG Generation in Complex Document and Meeting Scenarios

Yiming Xu, Jihua Kang, Chunsai Du +5

In demanding professional environments and meeting review scenarios, lengthy text often imposes a high cognitive load. To facilitate efficient information communication, transformi…

cs.CV2026

From Local Matches to Global Masks: Template-Guided Instance Detection and Segmentation in Open-World Scenes

Qifan Zhang, Sai Haneesh Allu, Jikai Wang +2

Detecting and segmenting novel object instances in open-world environments is a fundamental problem in robotic perception. Given only a small set of template images, a robot must l…

cs.CV2025

Continual Distillation Learning for Rehearsal-Free Class-Incremental Learning via Decoupled Prompting

Qifan Zhang, Yunhui Guo, Yu Xiang

Prompt-based continual learning has shown strong performance in rehearsal-free class-incremental learning by adapting learnable prompts while freezing a pre-trained Vision Transfor…

cs.AI2025

R: A LLM Based Novel-to-Screenplay Generation Framework with Causal Plot Graphs

Zefeng Lin, Yi Xiao, Zhiqiang Mo +8

Automatically adapting novels into screenplays is important for the TV, film, or opera industries to promote products with low costs. The strong performances of large language mode…

cs.CV2025

HO-Cap: A Capture System and Dataset for 3D Reconstruction and Pose Tracking of Hand-Object Interaction

Jikai Wang, Qifan Zhang, Yu-Wei Chao +3

We introduce a data capture system and a new dataset, HO-Cap, for 3D reconstruction and pose tracking of hands and objects in videos. The system leverages multiple RGBD cameras and…

cs.CV2024

CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities

Rohith Peddi, Shivvrat Arya, Bharath Challa +10

Following step-by-step procedures is an essential component of various activities carried out by individuals in their daily lives. These procedures serve as a guiding framework tha…