papers

Publications (8)

cs.DC2026

Omni-Flow: A Unified Workflow Orchestration and Distributed KV Cache Sharing Framework for Multimodal Inference

Bin Xiao, Jingfu Dong, Changran Wang +5

As large language model (LLM) inference evolves from text-only to multimodal paradigms, inference systems face three challenges: (1) flexible orchestration of multimodal workflows,…

cs.CV2025

TARA: Token-Aware LoRA for Composable Personalization in Diffusion Models

Yuqi Peng, Lingtao Zheng, Yufeng Yang +4

Personalized text-to-image generation aims to synthesize novel images of a specific subject or style using only a few reference images. Recent methods based on Low-Rank Adaptation…

cs.CV2026

ControlLight: Towards Controllable, Consistent, and Generalizable Low-Light Enhancement

Yufeng Yang, Jianzhuang Liu, Jisheng Chu +4

Existing deep learning-based low-light enhancement methods are typically trained on limited datasets with single enhancement targets, which restricts their generalization ability a…

cs.CV2026

LongCat-Next: Lexicalizing Modalities as Discrete Tokens

Meituan LongCat Team, Bin Xiao, Chao Wang +86

The prevailing Next-Token Prediction (NTP) paradigm has driven the success of large language models through discrete autoregressive modeling. However, contemporary multimodal syste…

cs.CV2025

GLAD: Generalizable Tuning for Vision-Language Models

Yuqi Peng, Pengfei Wang, Jianzhuang Liu +1

Pre-trained vision-language models, such as CLIP, show impressive zero-shot recognition ability and can be easily transferred to specific downstream tasks via prompt tuning, even w…

cs.MM2025

LongCat-Flash-Omni Technical Report

Meituan LongCat Team, Bairui Wang, Bayan +129

We introduce LongCat-Flash-Omni, a state-of-the-art open-source omni-modal model with 560 billion parameters, excelling at real-time audio-visual interaction. By adopting a curricu…