RAGtrap: Source Revocation and Indexed Provenance Lookup for Poisoned RAG Corpora

Resumo


Retrieval-augmented generation (RAG) answers questions from passages retrieved out of a corpus, and poisoned passages can steer those answers: in a published attack, 5 passages per target question sufficed to produce the attacker’s chosen answer. Once a source is known to be compromised, recovery means finding the passages it supplied and removing them without discarding unrelated content. RAGTRAP records a signed provenance entry for every passage at ingestion, indexed by source and by content hash. Tracing a suspect passage is then one hash lookup and no language-model request, against one or more per passage for post-incident attribution in the literature. Revoking a source removes only its passages, whereas deleting whole documents also discards benign ones. Exact hashing cannot attribute content changed after ingestion, nor content supplied by more than one source, so RAGTRAP supports recovery from a known compromised source but does not detect or prevent poisoning.

Referências

Agarwal, S. and Sen, A. (2026). The impact of Google AI Overviews on publisher traffic and user experience: Evidence from a field experiment. SSRN Working Paper 6513059.

Andersen, T. E., Avalos, A. M., Dagher, G. G., and Long, M. (2025). D-RAG: A privacy-preserving framework for decentralized RAG using blockchain. In Computer Science & Information Technology (CS & IT), pages 183–198.

Broder, A. Z. (1997). On the resemblance and containment of documents. In Compression and Complexity of Sequences (SEQUENCES).

Carlini, N. et al. (2024). Poisoning web-scale training datasets is practical. In IEEE S&P. Charikar, M. S. (2002). Similarity estimation techniques from rounding algorithms. In STOC.

Huang, Y. and Huang, J. X. (2026). A survey on retrieval-augmented text generation for large language models. ACM Computing Surveys, 58(12).

Indyk, P. and Motwani, R. (1998). Approximate nearest neighbors: Towards removing the curse of dimensionality. In STOC.

Kim, S., Kim, J., Jeon, Y., and Lee, G. G. (2025). Safeguarding RAG pipelines with GMTP: A gradient-based masked token probability method for poisoned document detection. In ACL Findings.

Kwiatkowski, T. et al. (2019). Natural Questions: A benchmark for question answering research. TACL, 7.

Namazi, M., Nemecek, A., and Ayday, E. (2025). ZKPROV: A zero-knowledge approach to dataset provenance for large language models. arXiv:2506.20915.

Newman, Z., Meyers, J. S., and Torres-Arias, S. (2022). Sigstore: Software signing for everybody. In ACM CCS, pages 2353–2367.

Ono, T. (2026). MerkleSpeech: Public-key verifiable, chunk-localised speech provenance via perceptual fingerprints and merkle commitments. arXiv:2602.10166.

Pathmanathan, P., Panaitescu-Liess, M.-A., Chiang, C.-Y. J., and Huang, F. (2025). RAGPart & RAGMask: Retrieval-stage defenses against corpus poisoning in retrieval-augmented generation. arXiv:2512.24268.

Patil, K. (2026). RAGShield: Provenance-verified defense-in-depth against knowledge base poisoning in government retrieval-augmented generation systems. arXiv:2604.00387v1. Cited v1 (1 Apr 2026); v2 title and claims differ.

Ray, J.-M. L. (2025). Policy-governed RAG: Research design study. arXiv:2510.19877. Shan, S., Bhagoji, A. N., Zheng, H., and Zhao, B. Y. (2022). Poison forensics: Traceback of data poisoning attacks in neural networks. In USENIX Security.

Tan, X. et al. (2025). RevPRAG: Revealing poisoning attacks in retrieval-augmented generation through LLM activation analysis. In EMNLP Findings.

Thakur, N. et al. (2021). BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In NeurIPS Datasets and Benchmarks.

Wang, L. et al. (2022). Text embeddings by weakly-supervised contrastive pre-training. arXiv:2212.03533.

Wanger, J. (2026). VectorSmuggle: Steganographic exfiltration in embedding stores and a cryptographic provenance defense. arXiv:2605.13764.

Xu, Y. et al. (2026). Securing retrieval-augmented generation: A taxonomy of attacks, defenses, and future directions. arXiv:2604.08304.

Zhang, B. et al. (2025). Traceback of poisoning attacks to retrieval-augmented generation. In WWW.

Zhang, B. et al. (2026). Who taught the lie? responsibility attribution for poisoned knowledge in retrieval-augmented generation. In IEEE S&P. arXiv:2509.13772.

Zou, W., Geng, R., Wang, B., and Jia, J. (2025). PoisonedRAG: Knowledge corruption attacks to retrieval-augmented generation of large language models. In USENIX Security.
Publicado
01/09/2026
KAPELINSKI, Cristhian; KREUTZ, Diego. RAGtrap: Source Revocation and Indexed Provenance Lookup for Poisoned RAG Corpora. In: WORKSHOP DE TRABALHOS DE INICIAÇÃO CIENTÍFICA E DE GRADUAÇÃO - SIMPÓSIO BRASILEIRO DE CIBERSEGURANÇA (SBSEG), 26. , 2026, Armação dos Búzios/RJ. Anais [...]. Porto Alegre: Sociedade Brasileira de Computação, 2026 . p. 687-700. DOI: https://doi.org/10.5753/sbseg_estendido.2026.29850.

Artigos mais lidos do(s) mesmo(s) autor(es)

<< < 1 2 3 4 5 6 7 8 9 10 > >>