16 papers
MoVA: Learning Asymmetric Dual Projections for Modular Long Video-Text Alignment
Peiyuan Zhu, Shaoan Xie, Zijian Li +5
Contrastive pre-training has propelled video-text alignment, yet models often inherit the critical limitations of their image-text predecessors like CLIP, resulting in entangled re…
Beyond Perplexity: A Behavioral Evaluation Framework for Deployment-Memory Claims in LLM Test-Time Training
Xiangchen Song, Zhenhao Chen, Lingjing Kong +4
Large language model test-time training (TTT) is often evaluated through local proxy metrics: models are updated on recent tokens, retrieved context, target-domain data, or verifia…
A Dialogue between Causal and Traditional Representation Learning: Toward Mutual Benefits in a Unified Formulation
Yan Li, Yuewen Sun, Shaoan Xie +4
Causal representation learning (CRL) and traditional representation learning have largely developed along different trajectories. Traditional representation learning has been drive…
SEDGE: Structural Extrapolated Data Generation
Kun Zhang, Jiaqi Sun, Yiqing Li +3
This paper aims to address the challenge of data generation beyond the training data and proposes a framework for Structural Extrapolated Data GEneration (SEDGE) based on suitable…
From Generalist to Specialist Representation
Yujia Zheng, Fan Feng, Yuke Li +3
Given a generalist model, learning a task-relevant specialist representation is fundamental for downstream applications. Identifiability, the asymptotic guarantee of recovering the…
The Power of Order: Fooling LLMs with Adversarial Table Permutations
Xinshuai Dong, Haifeng Chen, Xuyuan Liu +5
Large Language Models have achieved remarkable success and are increasingly deployed in critical applications involving tabular data, such as Table Question Answering. However, the…