activity
20222024
most citedMIT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

14 citations · 30 across the 12 of their papers we have counts for

collaborators

12 papers

eess.AS2024

Dynamic Encoder Size Based on Data-Driven Layer-wise Pruning for Speech Recognition

Jingjing Xu, Wei Zhou, Zijian Yang +2

Varying-size models are often required to deploy ASR systems under different hardware and/or application constraints such as memory and latency. To avoid redundant training and opt…

cs.CL2024

InfiAgent-DABench: Evaluating Agents on Data Analysis Tasks

Xueyu Hu, Ziyu Zhao, Shuang Wei +14

In this paper, we introduce InfiAgent-DABench, the first benchmark specifically designed to evaluate LLM-based agents on data analysis tasks. These tasks require agents to end-to-e…

cs.CL20238 cited

Extrapolating Large Language Models to Non-English by Aligning Languages

Wenhao Zhu, Yunzhe Lv, Qingxiu Dong +6

Existing large language models show disparate capability across different languages, due to the imbalance in the training data. Their performances on English tasks are often strong…

cs.CL2023

INK: Injecting kNN Knowledge in Nearest Neighbor Machine Translation

Wenhao Zhu, Jingjing Xu, Shujian Huang +2

Neural machine translation has achieved promising results on many translation tasks. However, previous studies have shown that neural models induce a non-smooth representation spac…

cs.CV202314 cited

MIT: A Large-Scale Dataset towards Multi-Modal Multilingual Instruction Tuning

Lei Li, Yuwei Yin, Shicheng Li +9

Instruction tuning has significantly advanced large language models (LLMs) such as ChatGPT, enabling them to align with human instructions across diverse tasks. However, progress i…

cs.CL20232 cited

Can Language Models Understand Physical Concepts?

Lei Li, Jingjing Xu, Qingxiu Dong +4

Language models~(LMs) gradually become general-purpose interfaces in the interactive and embodied world, where the understanding of physical concepts is an essential prerequisite.…