activity
20212026
most citedVision-Language Pre-Training with Triple Contrastive Learning

14 citations · 33 across the 5 of their papers we have counts for

collaborators

7 papers

cs.SE2026

TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents

Shoufa Chen, Luyuan Wang, Xuan Yang +7

As large language models and harness frameworks continue to advance, agents operating in terminals are increasingly capable of performing a broader range of general computer-use ta…

cs.CL2025

Learning LLM Preference over Intra-Dialogue Pairs: A Framework for Utterance-level Understandings

Xuanqing Liu, Luyang Kong, Wei Niu +6

Large language models (LLMs) have demonstrated remarkable capabilities in handling complex dialogue tasks without requiring use case-specific fine-tuning. However, analyzing live d…

cs.CV20225 cited

Multi-modal Alignment using Representation Codebook

Jiali Duan, Liqun Chen, Son Tran +4

Aligning signals from different modalities is an important step in vision-language representation learning as it affects the performance of later stages such as cross-modality fusi…

cs.CV202214 cited

Vision-Language Pre-Training with Triple Contrastive Learning

Jinyu Yang, Jiali Duan, Son Tran +6

Vision-language representation learning largely benefits from image-text alignment through contrastive losses (e.g., InfoNCE loss). The success of this alignment strategy is attrib…

cs.CL20215 cited

Magic Pyramid: Accelerating Inference with Early Exiting and Token Pruning

Xuanli He, Iman Keivanloo, Yi Xu +4

Pre-training and then fine-tuning large language models is commonly used to achieve state-of-the-art performance in natural language processing (NLP) tasks. However, most pre-train…

cs.CV2021

MLIM: Vision-and-Language Model Pre-training with Masked Language and Image Modeling

Tarik Arici, Mehmet Saygin Seyfioglu, Tal Neiman +5

Vision-and-Language Pre-training (VLP) improves model performance for downstream tasks that require image and text inputs. Current VLP approaches differ on (i) model architecture (…