33 citations · 43 across the 17 of their papers we have counts for
8 papers · 1 filter
Evolving Subnetwork Training for Large Language Models
Hanqi Li, Lu Chen, Da Ma +3
Large language models have ushered in a new era of artificial intelligence research. However, their substantial training costs hinder further development and widespread adoption. I…
Sparsity-Accelerated Training for Large Language Models
Da Ma, Lu Chen, Pengyu Wang +6
Large language models (LLMs) have demonstrated proficiency across various natural language processing (NLP) tasks but often require additional training, such as continual pre-train…
Hierarchical Multimodal Pre-training for Visually Rich Webpage Understanding
Hongshen Xu, Lu Chen, Zihan Zhao +4
The growing prevalence of visually rich documents, such as webpages and scanned/digital-born documents (images, PDFs, etc.), has led to increased interest in automatic document und…
A BiRGAT Model for Multi-intent Spoken Language Understanding with Hierarchical Semantic Frames
Hongshen Xu, Ruisheng Cao, Su Zhu +4
Previous work on spoken language understanding (SLU) mainly focuses on single-intent settings, where each input utterance merely contains one user intent. This configuration signif…
ASTormer: An AST Structure-aware Transformer Decoder for Text-to-SQL
Ruisheng Cao, Hanchong Zhang, Hongshen Xu +4
Text-to-SQL aims to generate an executable SQL program given the user utterance and the corresponding database schema. To ensure the well-formedness of output SQLs, one prominent a…
TeCS: A Dataset and Benchmark for Tense Consistency of Machine Translation
Yiming Ai, Zhiwei He, Kai Yu +1
Tense inconsistency frequently occurs in machine translation. However, there are few criteria to assess the model's mastery of tense prediction from a linguistic perspective. In th…