◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Xiaoya Lu

6 papers here

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2
  • middle author3

Across the 5 of 6 papers where every author was matched, so the position is known.

fields
  • cs.AI3
  • cs.CL1
  • cs.CR1
  • cs.MA1

identity via Semantic Scholar / OpenAlex

collaborators
Showing cs.CLShow all

2 papers · 1 filter

cs.CL2025

LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions

Xuhao Hu, Peng Wang, Xiaoya Lu +3

Previous research has shown that LLMs finetuned on malicious or incorrect completions within narrow domains (e.g., insecure code or incorrect medical advice) can become broadly mis…

cs.CL2024

LLMs know their vulnerabilities: Uncover Safety Gaps through Natural Distribution Shifts

Qibing Ren, Hao Li, Dongrui Liu +7

Safety concerns in large language models (LLMs) have gained significant attention due to their exposure to potentially harmful data during pre-training. In this paper, we identify…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.