5 papers
LEMON: How Well Do MLLMs Perform Temporal Multimodal Understanding on Instructional Videos?
Zhuang Yu, Lei Shen, Jing Zhao +1
Recent multimodal large language models (MLLMs) have shown remarkable progress across vision, audio, and language tasks, yet their performance on long-form, knowledge-intensive, an…
BARD: budget-aware reasoning distillation
Lujie Niu, Lei Shen, Yi Jiang +4
While long Chain-of-Thought (CoT) distillation effectively transfers reasoning capability to smaller language models, the reasoning process often remains redundant and computationa…
QAgent: A modular Search Agent with Interactive Query Understanding
Yi Jiang, Lei Shen, Lujie Niu +3
Large language models (LLMs) excel at natural language tasks but are limited by their static parametric knowledge, especially in knowledge-intensive task. Retrieval-augmented gener…
TDR: Task-Decoupled Retrieval with Fine-Grained LLM Feedback for In-Context Learning
Yifu Chen, Bingchen Huang, Zhiling Wang +4
In-context learning (ICL) has become a classic approach for enabling LLMs to handle various tasks based on a few input-output examples. The effectiveness of ICL heavily relies on t…
SEO: Stochastic Experience Optimization for Large Language Models
Jitao Xu, Hongyun Zhou, Lei Shen +3
Large Language Models (LLMs) can benefit from useful experiences to improve their performance on specific tasks. However, finding helpful experiences for different LLMs is not obvi…