6 papers
When Correct Demonstrations Hurt: Rethinking the Role of Exemplars in In-Context Learning
Chenghao Qiu, Chunli Peng, Yufeng Yang +2
In-context learning (ICL) is often motivated by the intuition that demonstrations help because they provide correct input-output examples. However, we reveal a counterintuitive phe…
Seeing Together: Multi-Robot Cooperative Egocentric Spatial Reasoning with Multimodal Large Language Models
Kunyu Peng, Zhikun Zhou, Kailun Yang +9
Multimodal Large Language Models (MLLMs) have made substantial progress in egocentric video understanding, but their ability to reason cooperatively from multiple embodied viewpoin…
CoCo: Code as CoT for Text-to-Image Preview and Rare Concept Generation
Haodong Li, Chunmei Qing, Huanyu Zhang +11
Recent advancements in Unified Multimodal Models (UMMs) have significantly advanced text-to-image (T2I) generation, particularly through the integration of Chain-of-Thought (CoT) r…
Entropy-Aware On-Policy Distillation of Language Models
Woogyeol Jin, Taywon Min, Yongjin Yang +5
On-policy distillation is a promising approach for transferring knowledge between language models, where a student learns from dense token-level signals along its own trajectories.…
EDTC: enhance depth of text comprehension in automated audio captioning
Liwen Tan, Yin Cao, Yi Zhou
Modality discrepancies have perpetually posed significant challenges within the realm of Automated Audio Captioning (AAC) and across all multi-modal domains. Facilitating models in…
Balanced SNR-Aware Distillation for Guided Text-to-Audio Generation
Bingzhi Liu, Yin Cao, Haohe Liu +1
Diffusion models have demonstrated promising results in text-to-audio generation tasks. However, their practical usability is hindered by slow sampling speeds, limiting their appli…