14 papers
SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model
Zhennan Chen, Tianxing Shi, Pengcheng Xu +5
VAR has gained widespread popularity due to its next-scale prediction paradigm. However, it faces substantial performance bottlenecks when handling complex scenes with multiple obj…
Quo Vadis, World Modeling?
Yu Yang, Xuemeng Yang, Licheng Wen +17
Continually improving agents require dynamic interaction feedback beyond static supervision, yet direct real-environment interaction is costly, slow, unsafe, and hard to paralleliz…
RoMeRL: Balancing Feedback Coverage and the Memory-Reward Trap in Self-Evolving Agent Memory via Reduced-Order Utility States
Yi Yang, Zhennan Chen, Yihong Zhuang +5
Learning-based memory systems for self-evolving LLM agents face two tightly coupled challenges. First, trajectory-indexed utilities grow with the interaction history, thereby dispe…
Spiking Pyramid Wavelet Transformation for High-efficient and Low-energy Image Restoration
Chen Zhao, Xiantao Hu, Song Wu +5
Spiking neural networks (SNNs) have garnered significant interest in computer vision due to their potential for efficiency and biological inspiration. While spiking CNN-based metho…
LUVE : Latent-Cascaded Ultra-High-Resolution Video Generation with Dual Frequency Experts
Chen Zhao, Jiawei Chen, Hongyu Li +6
Recent advances in video diffusion models have significantly improved visual quality, yet ultra-high-resolution (UHR) video generation remains a formidable challenge due to the com…
VINS-120K: Ultra High-Resolution Image Editing with A Large-Scale Dataset
Zhizhou Chen, Shanyan Guan, Zhanxin Gao +6
Directly editing ultra-high-resolution (UHR) images is valuable but underexplored, primarily due to the lack of high-quality data and the challenge in modeling high-frequency textu…