Score: 0

Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks

Published: January 6, 2026 | arXiv ID: 2601.03420v1

By: Zhakshylyk Nurlanov, Frank R. Schmidt, Florian Bernard

Potential Business Impact:

Breaks AI safety rules without needing special computer knowledge.

Business Areas:

Natural Language Processing Artificial Intelligence, Data and Analytics, Software

As Large Language Models (LLMs) are increasingly deployed in safety-critical domains, rigorously evaluating their robustness against adversarial jailbreaks is essential. However, current safety evaluations often overestimate robustness because existing automated attacks are limited by restrictive assumptions. They typically rely on handcrafted priors or require white-box access for gradient propagation. We challenge these constraints by demonstrating that token-level iterative optimization can succeed without gradients or priors. We introduce RAILS (RAndom Iterative Local Search), a framework that operates solely on model logits. RAILS matches the effectiveness of gradient-based methods through two key innovations: a novel auto-regressive loss that enforces exact prefix matching, and a history-based selection strategy that bridges the gap between the proxy optimization objective and the true attack success rate. Crucially, by eliminating gradient dependency, RAILS enables cross-tokenizer ensemble attacks. This allows for the discovery of shared adversarial patterns that generalize across disjoint vocabularies, significantly enhancing transferability to closed-source systems. Empirically, RAILS achieves near 100% success rates on multiple open-source models and high black-box attack transferability to closed-source systems like GPT and Gemini.

Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent

Machine Learning (CS)

Stops smart computers from being tricked.

20 Aug 2025 2

91%

Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints

Machine Learning (CS)

Makes AI more easily tricked into bad behavior.

25 Feb 2025 1

90%

Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations

Cryptography and Security

Stops AI from saying bad or unsafe things.

24 Nov 2025 1

View PDF Login to Bookmark

Page Count

17 pages

Jailbreaking LLMs Without Gradients or Priors: Effective and Transferable Attacks

Breaks AI safety rules without needing special computer knowledge.

Technical Abstract

Universal and Transferable Adversarial Attack on Large Language Models Using Exponentiated Gradient Descent

Guiding not Forcing: Enhancing the Transferability of Jailbreaking Attacks on LLMs via Removing Superfluous Constraints

Defending Large Language Models Against Jailbreak Exploits with Responsible AI Considerations