2 papers
cs.CL2022
Syntax-guided Localized Self-attention by Constituency Syntactic Distance
Shengyuan Hou, Jushi Kai, Haotian Xue +5
Recent works have revealed that Transformers are implicitly learning the syntactic information in its lower layers from data, albeit is highly dependent on the quality and scale of…
cs.LG2021
Counterfactual Adversarial Learning with Representation Interpolation
Wei Wang, Boxin Wang, Ning Shi +4
Deep learning models exhibit a preference for statistical fitting over logical reasoning. Spurious correlations might be memorized when there exists statistical bias in training da…