most citedThe Uncertainty-based Retrieval Framework for Ancient Chinese CWS and POS

5 citations · 5 across the 5 of their papers we have counts for

collaborators

8 papers

cs.CL2024

LongSafety: Enhance Safety for Long-Context LLMs

Mianqiu Huang, Xiaoran Liu, Shaojun Zhou +11

Recent advancements in model architectures and length extrapolation techniques have significantly extended the context length of large language models (LLMs), paving the way for th…

cs.CL2024

Case2Code: Scalable Synthetic Data for Code Generation

Yunfan Shao, Linyang Li, Yichuan Ma +11

Large Language Models (LLMs) have shown outstanding breakthroughs in code generation. Recent work improves code LLMs by training on synthetic data generated by some powerful LLMs,…

cs.CL2024

DenoSent: A Denoising Objective for Self-Supervised Sentence Representation Learning

Xinghao Wang, Junliang He, Pengyu Wang +3

Contrastive-learning-based methods have dominated sentence representation learning. These methods regularize the representation space by pulling similar sentence representations cl…

cs.CL2024

InferAligner: Inference-Time Alignment for Harmlessness through Cross-Model Guidance

Pengyu Wang, Dong Zhang, Linyang Li +5

With the rapid development of large language models (LLMs), they are not only used as general-purpose AI assistants but are also customized through further fine-tuning to meet the…

cs.CL2023

Watermarking LLMs with Weight Quantization

Linyang Li, Botian Jiang, Pengyu Wang +3

Abuse of large language models reveals high risks as large language models are being deployed at an astonishing speed. It is important to protect the model weights to avoid malicio…

cs.CL2023

PerturbScore: Connecting Discrete and Continuous Perturbations in NLP

Linyang Li, Ke Ren, Yunfan Shao +2

With the rapid development of neural network applications in NLP, model robustness problem is gaining more attention. Different from computer vision, the discrete nature of texts m…