collaborators

6 papers

cs.LG2025

DUMP: Automated Distribution-Level Curriculum Learning for RL-based LLM Post-training

Zhenting Wang, Guofeng Cui, Yu-Jhe Li +2

Recent advances in reinforcement learning (RL)-based post-training have led to notable improvements in large language models (LLMs), particularly in enhancing their reasoning capab…

cs.CV2025

CalibQuant: 1-Bit KV Cache Quantization for Multimodal LLMs

Insu Han, Zeliang Zhang, Zhiyuan Wang +8

Multimodal Large Language Models (MLLMs) have demonstrated remarkable performance across diverse applications. However, their computational overhead during deployment remains a cri…

cs.CV2024

Prompt-Guided Mask Proposal for Two-Stage Open-Vocabulary Segmentation

Yu-Jhe Li, Xinyang Zhang, Kun Wan +3

We tackle the challenge of open-vocabulary segmentation, where we need to identify objects from a wide range of categories in different environments, using text prompts as our inpu…

cs.CV2024

Treat Visual Tokens as Text? But Your MLLM Only Needs Fewer Efforts to See

Zeliang Zhang, Phu Pham, Wentian Zhao +6

By treating visual tokens from visual encoders as text tokens, Multimodal Large Language Models (MLLMs) have achieved remarkable progress across diverse visual understanding tasks,…

cs.CL2024

Efficient Self-Improvement in Multimodal Large Language Models: A Model-Level Judge-Free Approach

Shijian Deng, Wentian Zhao, Yu-Jhe Li +4

Self-improvement in multimodal large language models (MLLMs) is crucial for enhancing their reliability and robustness. However, current methods often rely heavily on MLLMs themsel…

cs.CV2024

A Data Perspective on Enhanced Identity Preservation for Diffusion Personalization

Xingzhe He, Zhiwen Cao, Nicholas Kolkin +4

Large text-to-image models have revolutionized the ability to generate imagery using natural language. However, particularly unique or personal visual concepts, such as pets and fu…