Cross-Lingual Indirect Prompt Injection Across Retrieval, Reranking, And Generation In Multilingual RAG
Abstract
External evidence can make retrieval-augmented generation (RAG) more informative, yet retrieved passages also provide a path for adversarial instructions to enter the model context. We examine that path in an English-Indonesian RAG system and track cross-lingual indirect prompt injection separately at retrieval, reranking, and generation. The experiment starts from 75 semantic items and evaluates every item in eight query/body/payload language combinations, for 600 paired trials. Qwen3-Embedding-0.6B and BGE-M3 produce top-20 candidate sets, BGE-reranker-v2-m3 reduces each set to five documents, and Qwen3-0.6B answers from the resulting context with either a standard prompt or an explicit trust-boundary prompt. Statistical uncertainty is estimated by resampling semantic items, and paired binary outcomes are modeled with generalized estimating equations. Poison documents reached the top 20 in 80.50% of Qwen trials and 37.67% of BGE-M3 trials (odds ratio 6.94, 95% CI 4.53-10.63). Top-five exposure was 18.50% and 16.17%, respectively. Standard end-to-end attack success was 6.33% for Qwen and 6.00% for BGE-M3; boundary-aware prompting lowered both rates to 1.33%, with no canary false positives. The results indicate that multilingual RAG security depends on several linked stages rather than generation alone.
Keywords
Downloads
References
[2] A. Asai, Z. Wu, Y. Wang, A. Sil, and H. Hajishirzi, “Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection,” in International Conference on Learning Representations, 2024.
[3] X. Zhang, N. Thakur, O. Ogundepo, E. Kamalloo, D. Alfonso-Hermelo, X. Li, Q. Liu, M. Rezagholizadeh, and J. Lin, “MIRACL: A Multilingual Retrieval Dataset Covering 18 Diverse Languages,” Transactions of the Association for Computational Linguistics, vol. 11, pp. 1114–1131, 2023, doi: 10.1162/tacl_a_00595.
[4] N. Muennighoff, N. Tazi, L. Magne, and N. Reimers, “MTEB: Massive Text Embedding Benchmark,” in Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pp. 2014–2037, 2023, doi: 10.18653/v1/2023.eacl-main.148.
[5] L. Ranaldi, B. Haddow, and A. Birch, “Multilingual Retrieval-Augmented Generation for Knowledge-Intensive Question Answering Task,” in Findings of the Association for Computational Linguistics: EACL 2026, pp. 697–716, 2026, doi: 10.18653/v1/2026.findings-eacl.35.
[6] W. Huo, X. Feng, B. Li, C. Fu, Y. Huang, H. Wang, and B. Qin, “Breaking Language Preference in Multilingual RAG via Language-Controllable Retrieval and Language-Agnostic Reasoning,” in Findings of the Association for Computational Linguistics: ACL 2026, pp. 7579–7589, 2026, doi: 10.18653/v1/2026.findings-acl.374.
[7] J. Chen, S. Xiao, P. Zhang, K. Luo, D. Lian, and Z. Liu, “M3-Embedding: Multi-Linguality, Multi-Functionality, Multi-Granularity Text Embeddings Through Self-Knowledge Distillation,” in Findings of the Association for Computational Linguistics: ACL 2024, pp. 2318–2335, 2024, doi: 10.18653/v1/2024.findings-acl.137.
[8] Y. Zhang, M. Li, D. Long, X. Zhang, H. Lin, B. Yang, P. Xie, A. Yang, D. Liu, J. Lin, F. Huang, and J. Zhou, “Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models,” arXiv preprint arXiv:2506.05176, 2025, doi: 10.48550/arXiv.2506.05176.
[9] S. Abdelnabi, K. Greshake, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” in Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec ’23), pp. 79–90, 2023, doi: 10.1145/3605764.3623985.
[10] J. Yi, Y. Xie, B. Zhu, E. Kiciman, G. Sun, X. Xie, and F. Wu, “Benchmarking and Defending against Indirect Prompt Injection Attacks on Large Language Models,” in KDD ’25: Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1, vol. 1, pp. 1809–1820, 2025, doi: 10.1145/3690624.3709179.
[11] Y. Liu, Y. Jia, R. Geng, J. Jia, and N. Z. Gong, “Formalizing and Benchmarking Prompt Injection Attacks and Defenses,” in 33rd USENIX Security Symposium (USENIX Security 24), pp. 1831–1847, 2024.
[12] Q. Zhan, Z. Liang, Z. Ying, and D. Kang, “InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents,” in Findings of the Association for Computational Linguistics: ACL 2024, pp. 10471–10506, 2024, doi: 10.18653/v1/2024.findings-acl.624.
[13] S. Chen, J. Piet, C. Sitawarin, and D. Wagner, “StruQ: Defending Against Prompt Injection with Structured Queries,” in 34th USENIX Security Symposium (USENIX Security 25), pp. 2383–2400, 2025.
[14] A. Shafran, R. Schuster, and V. Shmatikov, “Machine Against the RAG: Jamming Retrieval-Augmented Generation with Blocker Documents,” in 34th USENIX Security Symposium (USENIX Security 25), pp. 3787–3806, 2025.
[15] Z. Zhang, S. Li, Z. Zhang, X. Liu, H. Jiang, X. Tang, Y. Gao, Z. Li, H. Wang, Z. Tan, Y. Li, Q. Yin, B. Yin, and M. Jiang, “IHEval: Evaluating Language Models on Following the Instruction Hierarchy,” in Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies, vol. 1, pp. 8374–8398, 2025, doi: 10.18653/v1/2025.naacl-long.425.
[16] T. Wen, C. Wang, X. Yang, H. Tang, Y. Xie, L. Lyu, Z. Dou, and F. Wu, “Defending against Indirect Prompt Injection by Instruction Detection,” in Findings of the Association for Computational Linguistics: EMNLP 2025, pp. 19472–19487, 2025, doi: 10.18653/v1/2025.findings-emnlp.1060.
[17] Beijing Academy of Artificial Intelligence, “BAAI/bge-reranker-v2-m3 Model Card,” Hugging Face, 2024. Accessed: Aug. 19, 2026. [Online]. Available: https://huggingface.co/BAAI/bge-reranker-v2-m3
[18] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, et al., “Qwen3 Technical Report,” arXiv preprint arXiv:2505.09388, 2025, doi: 10.48550/arXiv.2505.09388.
