activity
20232026
collaborators

6 papers

cs.CV2026

EVA01: Unified Native 3D Understanding and Generation via Mixture-of-Transformers

Zongyuan Yang, Mingjing Yi, Wanli Ma +8

This paper addresses the challenge of integrating 3D meshes as a native modality within Multimodal Large Language Models (MLLMs). Diffusion-based large reconstruction models decoup…

cs.CV2025

Towards Scalable Training for Handwritten Mathematical Expression Recognition

Haoyang Li, Jiaqing Li, Jialun Cao +2

Large foundation models have achieved significant performance gains through scalable training on massive datasets. However, the field of \textbf{H}andwritten \textbf{M}athematical…

cs.GR2024

DirectL: Efficient Radiance Fields Rendering for 3D Light Field Displays

Zongyuan Yang, Baolin Liu, Yingde Song +4

Autostereoscopic display, despite decades of development, has not achieved extensive application, primarily due to the daunting challenge of 3D content creation for non-specialists…

cs.CV2023

DDG-Net: Discriminability-Driven Graph Network for Weakly-supervised Temporal Action Localization

Xiaojun Tang, Junsong Fan, Chuanchen Luo +3

Weakly-supervised temporal action localization (WTAL) is a practical yet challenging task. Due to large-scale datasets, most existing methods use a network pretrained in other data…

cs.CV2023

TextDiff: Mask-Guided Residual Diffusion Models for Scene Text Image Super-Resolution

Baolin Liu, Zongyuan Yang, Pengfei Wang +5

The goal of scene text image super-resolution is to reconstruct high-resolution text-line images from unrecognizable low-resolution inputs. The existing methods relying on the opti…

cs.CV2023

DocDiff: Document Enhancement via Residual Diffusion Models

Zongyuan Yang, Baolin Liu, Yongping Xiong +6

Removing degradation from document images not only improves their visual quality and readability, but also enhances the performance of numerous automated document analysis and reco…