◍wovepaper
SearchResearchersInstitutions
Sign in
researcher

Jinhwa Kim

2 papers hereh-index 355 citations5 works total

Matching runs newest-first, so older work may not be attached to this profile yet.

author position
  • first author2

Across the 2 of 2 papers where every author was matched, so the position is known.

fields
  • cs.CL1
  • cs.CR1

identity via Semantic Scholar / OpenAlex

collaborators

2 papers

cs.CR2025

Context Misleads LLMs: The Role of Context Filtering in Maintaining Safe Alignment of LLMs

Jinhwa Kim, Ian G. Harris

While Large Language Models (LLMs) have shown significant advancements in performance, various jailbreak attacks have posed growing safety and ethical risks. Malicious users often…

cs.CL2023

Robust Safety Classifier for Large Language Models: Adversarial Prompt Shield

Jinhwa Kim, Ali Derakhshan, Ian G. Harris

Large Language Models' safety remains a critical concern due to their vulnerability to adversarial attacks, which can prompt these systems to produce harmful responses. In the hear…

◍wovepaper

Papers, researchers and institutions, woven together.

Explore
  • Search
  • Researchers
  • Institutions
Account
  • Library
  • Chat
Data
  • arXiv.org
  • Semantic Scholar
  • OpenAlex
  • Latest RSS
AboutContactPrivacyDevelopersllms.txtopenapi.json
Not affiliated with arXiv. Researcher data from Semantic Scholar (ODC-BY) and OpenAlex.