activity
20232026
collaborators

6 papers

cs.LG2026

When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning

Chenghao Qiu, Chunli Peng, Yufeng Yang +2

In-context learning (ICL) is often motivated by the intuition that demonstrations help because they provide correct input-output examples. However, we reveal a counterintuitive phe…

cs.CV2026

Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models

Kunyu Peng, Zhikun Zhou, Kailun Yang +9

Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoin…

cs.AI2026

CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation

Haodong Li, Chunmei Qing, Huanyu Zhang +11

Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) r…

cs.LG2026

Entropy-Aware On-Policy Distillation of Language Models

Woogyeol Jin, Taywon Min, Yongjin Yang +5

On-policy distillation is a promising approach for transferring knowledge between language models, where a student learns from dense token-level signals along its own trajectories.…

cs.SD2024

EDTC: enhance depth of text comprehension in automated audio captioning

Liwen Tan, Yin Cao, Yi Zhou

Modality discrepancies have perpetually posed significant challenges within the realm of Automated Audio Captioning (AAC) and across all multi-modal domains. Facilitating models in…

cs.SD2023

Balanced SNR-Aware Distillation for Guided Text-to-Audio Generation

Bingzhi Liu, Yin Cao, Haohe Liu +1

Diffusion models have demonstrated promising results in text-to-audio generation tasks. However, their practical usability is hindered by slow sampling speeds, limiting their appli…