13 papers
SemLayer: Semantic-aware Generative Segmentation and Layer Construction for Abstract Icons
Haiyang Xu, Ronghuan Wu, Li-Yi Wei +5
Graphic icons are a cornerstone of modern design workflows, yet they are often distributed as flattened single-path or compound-path graphics, where the original semantic layering…
CyCLeGen: Cycle-Consistent Layout Prediction and Image Generation in Vision Foundation Models
Xiaojun Shan, Haoyu Shen, Yucheng Mao +5
We present CyCLeGen, a unified vision-language foundation model capable of both image understanding and image generation within a single autoregressive framework. Unlike existing v…
PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
Xiang Zhang, Sohyun Yoo, Hongrui Wu +3
We introduce PixARMesh, a method to autoregressively reconstruct complete 3D indoor scene meshes directly from a single RGB image. Unlike prior methods that rely on implicit signed…
VideoNSA: Native Sparse Attention Scales Video Understanding
Enxin Song, Wenhao Chai, Shusheng Yang +5
Video understanding in multimodal language models remains limited by context length: models often miss key transition frames and struggle to maintain coherence across long time sca…
Soft Tail-dropping for Adaptive Visual Tokenization
Zeyuan Chen, Kai Zhang, Zhuowen Tu +1
We present Soft Tail-dropping Adaptive Tokenizer (STAT), a 1D discrete visual tokenizer that adaptively chooses the number of output tokens per image according to its structural co…
CVP: Central-Peripheral Vision-Inspired Multimodal Model for Spatial Reasoning
Zeyuan Chen, Xiang Zhang, Haiyang Xu +2
We present a central-peripheral vision-inspired framework (CVP), a simple yet effective multimodal model for spatial reasoning that draws inspiration from the two types of human vi…