Taegyoon Kim
1 indexed paper
Recent (6 mo)
1With code
0Influential cites
0Benchmarked
0Publications per year
126
Top categories
NLP×1
Frequent co-authors
Research Timeline
2026
Can LLMs Reliably Self-Report Adversarial Prefills, and How?
This paper investigates how reliable large language models are in recognizing their own compromised outputs during adversarial prefill attacks in safety contexts.
Highlighted terms show continued research focus across papers