Publications (8)
Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference
Bin Xiao, Jingfu Dong, Changran Wang +5
As large language model (LLM) inference evolves from text-only to multimodal paradigms, inference systems face three challenges: (1) flexible orchestration of multimodal workflows,…
TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models
Yuqi Peng, Lingtao Zheng, Yufeng Yang +4
Personalized text-to-image generation aims to synthesize novel images of a specific subject or style using only a few reference images. Recent methods based on Low-Rank Adaptation…
ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement
Yufeng Yang, Jianzhuang Liu, Jisheng Chu +4
Existing deep learning-based low-light enhancement methods are typically trained on limited datasets with single enhancement targets, which restricts their generalization ability a…
LongCat-Next: Lexicalizing Modalities as Discrete Tokens
Meituan LongCat Team, Bin Xiao, Chao Wang +86
The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal syste…
GLAD: Generalizable Tuning for Vision-Language Models
Yuqi Peng, Pengfei Wang, Jianzhuang Liu +1
Pre-trained vision-language models, such as CLIP, show impressive zero-shot recognition ability and can be easily transferred to specific downstream tasks via prompt tuning, even w…
LongCat-Flash-Omni Technical Report
Meituan LongCat Team, Bairui Wang, Bayan +129
We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curricu…