4 papers
EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers
Zongyuan Yang, Mingjing Yi, Wanli Ma +8
This paper addresses the challenge of integrating 3D meshes as a native modality within Multimodal Large Language Models (MLLMs). Diffusion-based large reconstruction models decoup…
Towards Scalable Training for Handwritten Mathematical Expression Recognition
Haoyang Li, Jiaqing Li, Jialun Cao +2
Large foundation models have achieved significant performance gains through scalable training on massive datasets. However, the field of \textbf{H}andwritten \textbf{M}athematical…
TextDiff: Mask-Guided Residual Diffusion Models for Scene Text Image Super-Resolution
Baolin Liu, Zongyuan Yang, Pengfei Wang +5
The goal of scene text image super-resolution is to reconstruct high-resolution text-line images from unrecognizable low-resolution inputs. The existing methods relying on the opti…
DirectL: Efficient Radiance Fields Rendering for 3D Light Field Displays
Zongyuan Yang, Baolin Liu, Yingde Song +4
Autostereoscopic display, despite decades of development, has not achieved extensive application, primarily due to the daunting challenge of 3D content creation for non-specialists…