activity
20212024
most citedExtrapolating Large Language Models to Non-English by Aligning Languages

8 citations · 11 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CL2024

Multilingual Contrastive Decoding via Language-Agnostic Layers Skipping

Wenhao Zhu, Sizhe Liu, Shujian Huang +3

Decoding by contrasting layers (DoLa), is designed to improve the generation quality of large language models (LLMs) by contrasting the prediction probabilities between an early ex…

cs.CL2024

Getting More from Less: Large Language Models are Good Spontaneous Multilingual Learners

Shimao Zhang, Changjiang Gao, Wenhao Zhu +6

Recently, Large Language Models (LLMs) have shown impressive language capabilities. While most of the existing LLMs have very unbalanced performance across different languages, mul…

cs.CL20241 cited

MAPO: Advancing Multilingual Reasoning through Multilingual Alignment-as-Preference Optimization

Shuaijie She, Wei Zou, Shujian Huang +4

Though reasoning abilities are considered language-agnostic, existing LLMs exhibit inconsistent reasoning abilities across different languages, e.g., reasoning in the dominant lang…

cs.LG2023

PMNN:Physical Model-driven Neural Network for solving time-fractional differential equations

Zhiying Ma, Jie Hou, Wenhao Zhu +2

In this paper, an innovative Physical Model-driven Neural Network (PMNN) method is proposed to solve time-fractional differential equations. It establishes a temporal iteration sch…

cs.CV2023

Beyond Generic: Enhancing Image Captioning with Real-World Knowledge using Vision-Language Pre-Training Model

Kanzhi Cheng, Wenpo Song, Zheng Ma +3

Current captioning approaches tend to generate correct but "generic" descriptions that lack real-world knowledge, e.g., named entities and contextual information. Considering that…

cs.CL20238 cited

Extrapolating Large Language Models to Non-English by Aligning Languages

Wenhao Zhu, Yunzhe Lv, Qingxiu Dong +6

Existing large language models show disparate capability across different languages, due to the imbalance in the training data. Their performances on English tasks are often strong…