3 papers
cs.AI2026
PaDoc: Layout-Grounded Parallel Decoding for Document Parsing
Hao Yu, Jiabo Zhan, Kang Liu +8
End-to-end document parsers provide a unified interface, but serialize page layouts and regional contents into one autoregressive sequence. This formulation forces independent regi…
cs.CV2026
OmniAlpha: Aligning Transparency-Aware Generation via Multi-Task Unified Reinforcement Learning
Hao Yu, Jinglin Wang, Jiabo Zhan +7
Transparency-aware generation requires modeling not only RGB appearance but also alpha-based opacity and cross-layer composition, which are essential for tasks such as image mattin…
cs.CV2025
AlphaVAE: Unified End-to-End RGBA Image Reconstruction and Generation with Alpha-Aware Representation Learning
Zile Wang, Hao Yu, Jiabo Zhan +1
Recent advances in latent diffusion models have achieved remarkable results in high-fidelity RGB image synthesis by leveraging pretrained VAEs to compress and reconstruct pixel dat…