5 papers
HY3D-Bench: Generation of 3D Assets
Team Hunyuan3D, :, Bowen Zhang +22
While recent advances in neural representations and generative models have revolutionized 3D content creation, the field remains constrained by significant data processing bottlene…
An Adaptive Edge-Guided Dual-Network Framework for Fast QR Code Motion Deblurring
Jianping Li, Dongyang Guo, Wenjie Li +1
Unlike general image deblurring that prioritizes perceptual quality, QR code deblurring focuses on ensuring successful decoding. QR codes are characterized by highly structured pat…
ForCenNet: Foreground-Centric Network for Document Image Rectification
Peng Cai, Qiang Li, Kaicheng Yang +6
Document image rectification aims to eliminate geometric deformation in photographed documents to facilitate text recognition. However, existing methods often neglect the significa…
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
Fanheng Kong, Jingyuan Zhang, Hongzhi Zhang +7
Videos are unique in their integration of temporal elements, including camera, scene, action, and attribute, along with their dynamic relationships over time. However, existing ben…
Seed1.5-VL Technical Report
Dong Guo, Faming Wu, Feida Zhu +194
We present Seed1.5-VL, a vision-language foundation model designed to advance general-purpose multimodal understanding and reasoning. Seed1.5-VL is composed with a 532M-parameter v…