2 papers
cs.CV2025
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
Weili Xu, Enxin Song, Wenhao Chai +3
The challenge of long video understanding lies in its high computational complexity and prohibitive memory cost, since the memory and computation required by transformer-based LLMs…
cs.CV2025
An Empirical Study of GPT-4o Image Generation Capabilities
Sixiang Chen, Jinbin Bai, Zhuoran Zhao +16
The landscape of image generation has rapidly evolved, from early GAN-based approaches to diffusion models and, most recently, to unified generative architectures that seek to brid…