6 papers
Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference
Bin Xiao, Jingfu Dong, Changran Wang +5
As large language model (LLM) inference evolves from text-only to multimodal paradigms, inference systems face three challenges: (1) flexible orchestration of multimodal workflows,…
ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement
Yufeng Yang, Jianzhuang Liu, Jisheng Chu +4
Existing deep learning-based low-light enhancement methods are typically trained on limited datasets with single enhancement targets, which restricts their generalization ability a…
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Meituan LongCat Team, Bin Xiao, Chao Wang +86
The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal syste…
RealRestorer: Towards Generalizable Real-World Image Restoration with Large-Scale Image Editing Models
Yufeng Yang, Xianfang Zeng, Zhangqi Jiang +8
Image restoration under real-world degradations is critical for downstream tasks such as autonomous driving and object detection. However, existing restoration models are often lim…
TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
Yuqi Peng, Lingtao Zheng, Yufeng Yang +4
Personalized text-to-image generation aims to synthesize novel images of a specific subject or style using only a few reference images. Recent methods based on Low-Rank Adaptation…
GLAD: Generalizable Tuning for Vision-Language Models
Yuqi Peng, Pengfei Wang, Jianzhuang Liu +1
Pre-trained vision-language models, such as CLIP, show impressive zero-shot recognition ability and can be easily transferred to specific downstream tasks via prompt tuning, even w…