跳到主要导航 跳到搜索 跳到主要内容

Revealing the intrinsic ethical vulnerability of aligned large language models

  • Jiawei Lian
  • , Jianhong Pan
  • , Lefan Wang
  • , Yi Wang
  • , Xiaofei Wang
  • , Yingjie Lu
  • , Shaohui Mei
  • , Lap Pui Chau
  • Hong Kong Polytechnic University
  • Northwestern Polytechnical University Xian

科研成果: 期刊稿件文章同行评审

1 引用 (Scopus)

摘要

Large language models (LLMs) represent foundational advances toward artificial general intelligence, yet their alignment with human values via instruction tuning and preference learning achieves only superficial ethical compliance. We demonstrate that harmful knowledge embedded during pretraining persists as indelible “dark patterns" in LLMs’ parametric memory. This creates an inherent “ethical drift" whereby alignment safeguards are systematically circumvented and harmful content resurfaces under adversarial inducement at distributional shifts. Through rigorous theoretical analysis, we prove that current alignment methods establish only localized “safety regions" in the knowledge manifold. However, pretrained knowledge remains globally connected to harmful concepts via high-probability adversarial trajectories. We empirically validate these theoretical insights through a straightforward yet theoretically grounded methodology-semantic coherence inducement under distributional shifts. The effectiveness of this approach, achieving a 100% attack success rate across 22 out of 26 state-of-the-art aligned LLMs (including DeepSeek-R1, Llama-3, and Qwen3, among others), is not incidental but a direct consequence of our theoretical framework, demonstrating that the vulnerability is architectural rather than implementation-specific and revealing a fundamental structural weakness in current aligned LLMs.

源语言英语
文章编号4295
期刊Nature Communications
17
1
DOI
出版状态已出版 - 12月 2026

指纹

探究 'Revealing the intrinsic ethical vulnerability of aligned large language models' 的科研主题。它们共同构成独一无二的指纹。

引用此