16 papers
Video Generation Models as World Models: Efficient Paradigms, Architectures and Algorithms
Muyang He, Hanzhong Guo, Junxiong Lin +1
The rapid evolution of video generation has enabled models to simulate complex physical dynamics and long-horizon causalities, positioning them as potential world simulators. Howev…
Identifying Latent Concepts and Structures for Generalized Category Discovery
Boyang Dai, Chaoqi Chen, Yizhou Yu
Generalized Category Discovery (GCD) aims to recognize known classes while autonomously discovering novel ones in open-world settings. However, current approaches primarily focus o…
Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis
Boyang Dai, Chaoqi Chen, Yizhou Yu
Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models. Existing methods mostly focus on regular entangled representations to discrimin…
Leveraging Verifier-Based Reinforcement Learning in Image Editing
Hanzhong Guo, Jie Wu, Jie Liu +6
While Reinforcement Learning from Human Feedback (RLHF) has become a pivotal paradigm for text-to-image generation, its application to image editing remains largely unexplored. A k…
Parameters as Experts: Adapting Vision Models with Dynamic Parameter Routing
Meng Lou, Stanley Yu, Yizhou Yu
Adapting pre-trained vision models using parameter-efficient fine-tuning (PEFT) remains challenging, as it aims to achieve performance comparable to full fine-tuning using a minima…
Overcoming Catastrophic Forgetting in Visual Continual Learning with Reinforcement Fine-Tuning
Meng Lou, Hanzhong Guo, Linwei Chen +1
Recent studies suggest that Reinforcement Fine-Tuning (RFT) is inherently more resilient to catastrophic forgetting than Supervised Fine-Tuning (SFT). However, whether RFT (e.g., G…