collaborators

5 papers

cs.LG2026

MAG: MAnifold Guided Semi-Supervised Multi-modal In-Context Learning

Zirui Cheng, Xun Xu, Tiankai Chen +7

Few-shot in-context learning (ICL) with multi-modal large language models (MLLMs) enables task adaptation without parameter updates, but its performance is highly sensitive to the…

cs.CL2026

Simplex Relaxation for Discrete Diffusion

Jinya Sakurai, Patrick Pynadath, Satoshi Hayakawa +4

Discrete diffusion models for categorical generation are defined by a corruption kernel, which determines the intermediate state space and the associated reverse prediction problem…

cs.CL2026

MERaLiON-GR: Speech Gender Recognition Model for English and SEA Languages

Qiongqiong Wang, Ai Ti Aw, Nancy F. Chen +18

We present MERaLiON-GR, a speech gender recognition system that performs binary classification (female / male) on English and Southeast Asian (SEA) languages. The model finetunes M…

cs.SD2025

MERaLiON-SER: Robust Speech Emotion Recognition Model for English and SEA Languages

Hardik B. Sailor, Aw Ai Ti, Chen Fang Yih Nancy +26

We present MERaLiON-SER, a robust speech emotion recognition model designed for English and Southeast Asian languages. The model is trained using a hybrid objective combining weigh…

cs.CV2025

Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs

Yongyi Su, Haojie Zhang, Shijie Li +11

Multimodal large language models (MLLMs) have advanced rapidly in recent years. However, existing approaches for vision tasks often rely on indirect representations, such as genera…