papers

Publications (16)

cond-mat.mes-hall2010

Strain distributions in lattice-mismatched semiconductor core-shell nanowires

Niels Søndergaard, Yuhui He, Chun Fan +3

The authors study the elastic deformation field in lattice-mismatched core-shell nanowires with single and multiple shells. The authors consider infinite wires with a hexagonal cro…

cs.CL2020

Self-Explaining Structures Improve NLP Models

Zijun Sun, Chun Fan, Qinghong Han +4

Existing approaches to explaining deep learning models in NLP usually suffer from two major drawbacks: (1) the main model and the explaining model are decoupled: an additional prob…

cs.CV2026

UniFace: A Unified Fine-grained Face Understanding and Generation Model

Junzhe Li, Sifan Zhou, Liya Guo +9

Unified multimodal models (UMMs) have emerged as a powerful paradigm in fundamental cross-modality research, demonstrating significant potential in both image understanding and gen…

cs.CL2021

BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation Models

Kangjie Chen, Yuxian Meng, Xiaofei Sun +4

Pre-trained Natural Language Processing (NLP) models can be easily adapted to a variety of downstream language tasks. This significantly accelerates the development of language mod…

cs.CL2021

Folden: -Fold Ensemble for Out-Of-Distribution Detection

Xiaoya Li, Jiwei Li, Xiaofei Sun +5

Out-of-Distribution (OOD) detection is an important problem in natural language processing (NLP). In this work, we propose a simple yet effective framework Folden, which mimics…

cs.CL2020

Pair the Dots: Jointly Examining Training History and Test Stimuli for Model Interpretability

Yuxian Meng, Chun Fan, Zijun Sun +3

Any prediction from a model is made by a combination of learning history and test stimuli. This provides significant insights for improving model interpretability: {\it because of…

cs.AI2026

MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE

Junzhe Li, Yutao Cui, Tao Huang +8

Although GRPO substantially enhances flow matching models in human preference alignment of image generation, methods such as FlowGRPO and DanceGRPO still exhibit inefficiency due t…

cs.LG2021

A General Framework for Defending Against Backdoor Attacks via Influence Graph

Xiaofei Sun, Jiwei Li, Xiaoya Li +5

In this work, we propose a new and general framework to defend against backdoor attacks, inspired by the fact that attack triggers usually follow a \textsc{specific} type of attack…

stat.ML2021

Parameter Estimation for the SEIR Model Using Recurrent Nets

Chun Fan, Yuxian Meng, Xiaofei Sun +3

The standard way to estimate the parameters (e.g., the transmission rate ) of an SEIR model is to use grid search, where simulations are performed on each set…

cs.CL2020

Neural Semi-supervised Learning for Text Classification Under Large-Scale Pretraining

Zijun Sun, Chun Fan, Xiaofei Sun +3

The goal of semi-supervised learning is to utilize the unlabeled, in-domain dataset U to improve models trained on the labeled dataset D. Under the context of large-scale language-…

cs.CL2022

Summarize, Outline, and Elaborate: Long-Text Generation via Hierarchical Supervision from Extractive Summaries

Xiaofei Sun, Zijun Sun, Yuxian Meng +2

The difficulty of generating coherent long texts lies in the fact that existing models overwhelmingly focus on predicting local words, and cannot make high level plans on what to g…

cs.CL2021

Layer-wise Model Pruning based on Mutual Information

Chun Fan, Jiwei Li, Xiang Ao +3

The proposed pruning strategy offers merits over weight-based pruning techniques: (1) it avoids irregular memory access since representations and matrices can be squeezed into thei…

cs.CL2022

Dependency Parsing as MRC-based Span-Span Prediction

Leilei Gan, Yuxian Meng, Kun Kuang +4

Higher-order methods for dependency parsing can partially but not fully address the issue that edges in dependency trees should be constructed at the text span/subtree level rather…

cs.CL2022

Triggerless Backdoor Attack for NLP Tasks with Clean Labels

Leilei Gan, Jiwei Li, Tianwei Zhang +6

Backdoor attacks pose a new threat to NLP models. A standard strategy to construct poisoned data in backdoor attacks is to insert triggers (e.g., rare words) into selected sentence…

cs.CL2022

Sentence Similarity Based on Contexts

Xiaofei Sun, Yuxian Meng, Xiang Ao +4

Existing methods to measure sentence similarity are faced with two challenges: (1) labeled datasets are usually limited in size, making them insufficient to train supervised neural…

cs.CL2022

Paraphrase Generation as Unsupervised Machine Translation

Xiaofei Sun, Yufei Tian, Yuxian Meng +4

In this paper, we propose a new paradigm for paraphrase generation by treating the task as unsupervised machine translation (UMT) based on the assumption that there must be pairs o…