6 papers
Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
Haocheng Xi, Shuo Yang, Yilong Zhao +13
Despite rapid progress in autoregressive video diffusion, an emerging system algorithm bottleneck limits both deployability and generation capability: KV cache memory. In autoregre…
HiconAgent: History Context-aware Policy Optimization for GUI Agents
Xurui Zhou, Gongwei Chen, Yuquan Xie +6
Graphical User Interface (GUI) agents require effective use of historical context to perform sequential navigation tasks. While incorporating past actions and observations can impr…
StreamDiffusionV2: A Streaming System for Dynamic and Interactive Video Generation
Tianrui Feng, Zhi Li, Shuo Yang +11
Generative models are reshaping the live-streaming industry by redefining how content is created, styled, and delivered. Previous image-based streaming diffusion models have powere…
Jet-RL: Enabling On-Policy FP8 Reinforcement Learning with Unified Training and Rollout Precision Flow
Haocheng Xi, Charlie Ruan, Peiyuan Liao +7
Reinforcement learning (RL) is essential for enhancing the complex reasoning capabilities of large language models (LLMs). However, existing RL training pipelines are computational…
SparseTem: Boosting the Efficiency of CNN-Based Video Encoders by Exploiting Temporal Continuity
Kunyun Wang, Shuo Yang, Jieru Zhao +4
Deep learning models have become pivotal in the field of video processing and is increasingly critical in practical applications such as autonomous driving and object detection. Al…
A Survey of Multimodal Ophthalmic Diagnostics: From Task-Specific Approaches to Foundational Models
Xiaoling Luo, Ruli Zheng, Qiaojian Zheng +6
Visual impairment represents a major global health challenge, with multimodal imaging providing complementary information that is essential for accurate ophthalmic diagnosis. This…