Score: 2

Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation

Published: December 25, 2025 | arXiv ID: 2512.21681v1

By: Tian Li , Bo Lin , Shangwen Wang and more

Potential Business Impact:

Makes AI write unsafe computer code.

Business Areas:

Augmented Reality Hardware, Software

Retrieval-Augmented Code Generation (RACG) is increasingly adopted to enhance Large Language Models for software development, yet its security implications remain dangerously underexplored. This paper conducts the first systematic exploration of a critical and stealthy threat: backdoor attacks targeting the retriever component, which represents a significant supply-chain vulnerability. It is infeasible to assess this threat realistically, as existing attack methods are either too ineffective to pose a real danger or are easily detected by state-of-the-art defense mechanisms spanning both latent-space analysis and token-level inspection, which achieve consistently high detection rates. To overcome this barrier and enable a realistic analysis, we first developed VenomRACG, a new class of potent and stealthy attack that serves as a vehicle for our investigation. Its design makes poisoned samples statistically indistinguishable from benign code, allowing the attack to consistently maintain low detectability across all evaluated defense mechanisms. Armed with this capability, our exploration reveals a severe vulnerability: by injecting vulnerable code equivalent to only 0.05% of the entire knowledge base size, an attacker can successfully manipulate the backdoored retriever to rank the vulnerable code in its top-5 results in 51.29% of cases. This translates to severe downstream harm, causing models like GPT-4o to generate vulnerable code in over 40% of targeted scenarios, while leaving the system's general performance intact. Our findings establish that retriever backdooring is not a theoretical concern but a practical threat to the software development ecosystem that current defenses are blind to, highlighting the urgent need for robust security measures.

Give LLMs a Security Course: Securing Retrieval-Augmented Code Generation via Knowledge Injection

Cryptography and Security

Keeps computer code safe from bad instructions.

23 Apr 2025 0

91%

Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation

Cryptography and Security

Stops bad code from getting into computer programs.

5 Feb 2025 1

89%

Your RAG is Unfair: Exposing Fairness Vulnerabilities in Retrieval-Augmented Generation via Backdoor Attacks

Information Retrieval

Makes AI unfairly favor some people over others.

26 Sep 2025 0

View PDF Login to Bookmark

Country of Origin

🇨🇳 China

Repos / Data Links

github.com

Page Count

25 pages

Exploring the Security Threats of Retriever Backdoors in Retrieval-Augmented Code Generation

Makes AI write unsafe computer code.

Technical Abstract

Give LLMs a Security Course: Securing Retrieval-Augmented Code Generation via Knowledge Injection

Exploring the Security Threats of Knowledge Base Poisoning in Retrieval-Augmented Code Generation

Your RAG is Unfair: Exposing Fairness Vulnerabilities in Retrieval-Augmented Generation via Backdoor Attacks