activity
20192026
most citedXLM-T: Scaling up Multilingual Machine Translation with Pretrained Cross-lingual Transformer Encoders

23 citations · 68 across the 8 of their papers we have counts for

collaborators

12 papers

cs.AI2026

OmniDelta: Skill-Driven Budget Allocation for Token Compression in OmniLLMs

Haoyang Huang, Wenjie Huang, Tianqi Xu +14

Emerging Omni-modal Large Language Models (OmniLLMs) enable unified understanding of text, audio, and video, but their long audio-video token sequences introduce substantial memory…

cs.CL20246 cited

Synthetic Data (Almost) from Scratch: Generalized Instruction Tuning for Language Models

Haoran Li, Qingxiu Dong, Zhengyang Tang +17

We introduce Generalized Instruction Tuning (called GLAN), a general and scalable method for instruction tuning of Large Language Models (LLMs). Unlike prior work that relies on se…

cs.CL2024

Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models

Tianyi Tang, Wenyang Luo, Haoyang Huang +5

Large language models (LLMs) demonstrate remarkable multilingual capabilities without being pre-trained on specially curated multilingual parallel corpora. It remains a challenging…

cs.CL20222 cited

LVP-M3: Language-aware Visual Prompt for Multilingual Multimodal Machine Translation

Hongcheng Guo, Jiaheng Liu, Haoyang Huang +5

Multimodal Machine Translation (MMT) focuses on enhancing text-only translation with visual features, which has attracted considerable attention from both natural language processi…

cs.CV20221 cited

NÜWA-LIP: Language Guided Image Inpainting with Defect-free VQGAN

Minheng Ni, Chenfei Wu, Haoyang Huang +3

Language guided image inpainting aims to fill in the defective regions of an image under the guidance of text while keeping non-defective regions unchanged. However, the encoding p…

cs.CL20221 cited

SMDT: Selective Memory-Augmented Neural Document Translation

Xu Zhang, Jian Yang, Haoyang Huang +4

Existing document-level neural machine translation (NMT) models have sufficiently explored different context settings to provide guidance for target generation. However, little att…