Automated Red-Teaming Framework for RAG Systems
A seemingly harmless request for a project summary makes an AI assistant reveal a confidential code from an internal document. How can organizations detect such failures before deployment?
Gwerder, Oliver, 2026
Art der Arbeit Bachelor Thesis
Auftraggebende Trajecta GmbH
Betreuende Dozierende Bendel, Oliver
Views: 2
Retrieval-Augmented Generation (RAG) connects an AI assistant to company documents. This makes answers more useful, but it also places user requests, retrieved text, and internal information in the same processing context. A manipulated request can therefore turn legitimate access into unintended disclosure. Existing security tests often demonstrate individual attacks without showing whether a protection works repeatedly. The thesis examines how organizations can obtain comparable evidence in a synthetic Trajecta AG environment without using real personal or production data.
The project built an automated test workflow around a modular RAG system. It generates and reuses attack prompts, records responses, and checks whether synthetic test secrets become visible. Four matched conditions compare an unprotected system with prompt guarding, output masking, and both protections together. Broader hosted, multilingual, multi-turn, human-review, evaluator, and output-filter studies test transfer and measurement boundaries. Evidence sources remain separate so that broader coverage is not mistaken for one statistical sample.
Prompt guarding reduced configured-secret exposure, but the size and robustness of the effect varied across models and attack families. Output masking was highly reliable for registered secrets, while an unlisted control secret often remained visible; it therefore enforces a declared inventory rather than general confidentiality. The systems still completed all tested legitimate tasks in the primary matched conditions. Hosted transfer studies showed similar model-dependent effects and exposed a normalization boundary in deterministic masking. Evaluator studies showed that exact matching and semantic judging make different errors, supporting human review for ambiguous or high-impact cases. Across matched local and hosted short-horizon comparisons, adaptive variants did not show an aggregate advantage over fixed or direct strategies. The main result is therefore not a security certificate, but an auditable workflow that links attacks, evidence, uncertainty, review, and retesting so organizations can distinguish what was tested from what remains uncertain.
Studiengang: Business Artificial Intelligence (Bachelor)
Keywords Artificial Intelligence (AI), Retrieval-Augmented Generation (RAG), Cybersecurity, Automated Red Teaming, Prompt Injection, Data Privacy, Large Language Models (LLMs)
Vertraulichkeit: vertraulich