1 citations · 1 across the 1 of their papers we have counts for
6 papers
Enhancing Lexicon-Based Text Embeddings with Large Language Models
Yibin Lei, Tao Shen, Yu Cao +1
Recent large language models (LLMs) have demonstrated exceptional performance on general-purpose text embedding tasks. While dense embeddings have dominated related research, we in…
Seed2Scale: A Self-Evolving Data Engine for Embodied AI via Small to Large Model Synergy and Multimodal Evaluation
Cong Tai, Zhaoyu Zheng, Haixu Long +12
Existing data generation methods suffer from exploration limits, embodiment gaps, and low signal-to-noise ratios, leading to performance degradation during self-iteration. To addre…
ThinkQE: Query Expansion via an Evolving Thinking Process
Yibin Lei, Tao Shen, Andrew Yates
Effective query expansion for web search benefits from promoting both exploration and result diversity to capture multiple interpretations and facets of a query. While recent LLM-b…
MammothModa2: A Unified AR-Diffusion Framework for Multimodal Understanding and Generation
Tao Shen, Xin Wan, Taicai Chen +10
Unified multimodal models aim to integrate understanding and generation within a single framework, yet bridging the gap between discrete semantic reasoning and high-fidelity visual…
LLMs are Also Effective Embedding Models: An In-depth Overview
Chongyang Tao, Tao Shen, Shen Gao +6
Large language models (LLMs) have revolutionized natural language processing by achieving state-of-the-art performance across various tasks. Recently, their effectiveness as embedd…
A Survey on Knowledge Distillation of Large Language Models
Xiaohan Xu, Ming Li, Chongyang Tao +6
In the era of Large Language Models (LLMs), Knowledge Distillation (KD) emerges as a pivotal methodology for transferring advanced capabilities from leading proprietary LLMs, such…