A Review- XplainGuard: Dynamic Policy Enforcement in Cloud Infrastructure as Code: Integrating Generative AI for Auto- Remediation and Explainability
Main Article Content
Abstract
Infrastructure as Code (IaC) concentrates the security posture of an entire cloud estate into text files, where a single misconfigured attribute propagates a systemic vulnerability to every environment in which the configuration is applied. Established static analysis scanners detect such misconfigurations reliably but generate no corrective code, leaving remediation entirely manual. Recent generative approaches do produce repairs, yet three obstacles have prevented their unattended use. First, model responses are commonly returned as unstructured prose, so downstream tooling cannot parse them reliably. Second, generated patches are presented without any verification step, allowing syntactically plausible but invalid configuration code to reach the developer. Third, findings carry no security rationale, so practitioners cannot judge why a change is required or whether it is safe to apply. This paper presents XplainGuard, a command-line Terraform auditor that addresses all three. Model output is constrained to a strict four-field JSON schema and sanitized before parsing, making every response machine-readable. A self-verification guardrail then splices each generated patch back into a complete copy of the source file and executes terraform validate before the developer sees it; where validation fails, the exact diagnostic is returned to the model and a single bounded correction attempt is issued, which guarantees termination and caps cost at two model invocations per file. An explainability engine attaches a structured four-element rationale to every finding: the affected resource and attribute, the attack vector enabled, the compound-risk context arising from co-located resources, and the violated compliance control. Across 21 measured runs on a seven-file fixture, every response parsed successfully as valid JSON, and 14 of 15 generated patches (93.3%) were valid on first generation. The guardrail detected the single remaining invalid patch, which referenced an undeclared security group, and obtained a correct replacement, reducing the observed incorrect code rate from 6.7% to 0%. Detection coverage was 12 of 12 (100%) on files containing a single vulnerability and 10 of 12 (83.3%) on a file containing four simultaneously, where the explanation omitted one issue that the generated patch nonetheless corrected, a reporting limitation of the single-finding output schema rather than a detection failure. Independently executed scans show Trivy 0.72.0 failing to report two of four target vulnerability categories that Checkov 3.3.8 detects. The evaluation is deliberately small and controlled, and its scope and statistical limits are stated explicitly.
Keywords Infrastructure as Code, Terraform, dynamic policy enforcement, generative AI, large language models, auto-remediation, automated program repair, explainable AI, DevSecOps, cloud security, self-verification.
Article Details
Section
How to Cite
References
HashiCorp Inc., "Terraform: Infrastructure as Code Official Documentation," 2024. [Online]. Available: https://developer.hashicorp.com/terraform
Prisma Cloud / Bridgecrew, "Checkov: Policy-as-code for infrastructure," version 3.3.8, 2024. [Online]. Available: https://www.checkov.io
Aqua Security, "Trivy: Comprehensive security scanner," version 0.72.0, 2024. [Online]. Available: https://trivy.dev
G. De Vito, F. Palomba, and F. Ferrucci, "SecLLM: Enhancing Security Smell Detection in IaC with Large Language Models," IEEE Access, 2025, doi: 10.1109/ACCESS.2025.3637505.
E. Low, C. Cheh, and B. Chen, "Repairing Infrastructure-as-Code using Large Language Models," in Proc. 2024 IEEE Secure Development Conference (SecDev), Pittsburgh, PA, USA, 2024, pp. 20-27, doi: 10.1109/SecDev61143.2024.00008.
A. Rahman, C. Parnin, and L. Williams, "The Seven Sins: Security Smells in Infrastructure as Code Scripts," in Proc. IEEE/ACM 41st Int. Conf. on Software Engineering (ICSE), 2019, pp. 164-175.
P. Bhartiya, M. Bhatele, and A. A. Waoo, "Ensemble-Based Machine Learning Models for Real-Time Traffic Flow Prediction," Journal of Neonatal Surgery, vol. 14, no. 32s, pp. 6406-6419, 2025.
Reddy, D. B. L. L., Sumathi, Soni, L. N., M. R., & Nanthini, P. (2025). Deep Learning Algorithm For Sentiment Analysis Of E-Commerce. International Journal of Novel Research And Development (IJNRD), 10(5).
N. Saavedra and J. F. Ferreira, "GLITCH: Automated Polyglot Security Smell Detection in Infrastructure as Code," in Proc. 37th IEEE/ACM Int. Conf. on Automated Software Engineering (ASE), 2022, pp. 1-12.
N. Saavedra, J. Goncalves, M. Henriques, J. F. Ferreira, and A. Mendes, "InfraFix: Technology-Agnostic Repair of Infrastructure as Code," in Proc. 34th ACM SIGSOFT Int. Symp. on Software Testing and Analysis (ISSTA), Trondheim, Norway, 2025, doi: 10.1145/3713081.3731735.
Gautam, C. S., Soni, L. N., & Pandey, P. (2022). Clustering of Bigdata Using Genetic Algorithm in Hadoop Map Reduce. European Chemical Bulletin, 963–973.
H. Pearce, B. Tan, B. Ahmad, R. Karri, and B. Dolan-Gavitt, "Examining Zero-Shot Vulnerability Repair with Large Language Models," in Proc. IEEE Symp. on Security and Privacy (SP), 2023, pp. 2339-2356.
M. Jin, S. Shahriar, M. Tufano, X. Shi, S. Lu, N. Sundaresan, and A. Svyatkovskiy, "InferFix: End-to-End Program Repair with LLMs," in Proc. 31st ACM Joint European Software Engineering Conf. and Symp. on the Foundations of Software Engineering (ESEC/FSE), 2023, pp. 1646-1656.
C. Le Goues, M. Pradel, A. Roychoudhury, and S. Chandra, "Automatic Program Repair," IEEE Software, vol. 38, no. 4, pp. 22-27, Jul. 2021.
E. K. Smith, E. T. Barr, C. Le Goues, and Y. Brun, "Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair," in Proc. 10th Joint Meeting on Foundations of Software Engineering (ESEC/FSE), 2015, pp. 532-543.
A. Rahman, E. Farhana, C. Parnin, and L. Williams, "Gang of Eight: A Defect Taxonomy for Infrastructure as Code Scripts," in Proc. ACM/IEEE 42nd Int. Conf. on Software Engineering (ICSE), Seoul, Republic of Korea, 2020, pp. 752-764.
A. Rahman, M. R. Rahman, C. Parnin, and L. Williams, "Security Smells in Ansible and Chef Scripts: A Replication Study," ACM Trans. Softw. Eng. Methodol., vol. 30, no. 1, pp. 1-31, 2021.
M. Chiari, M. De Pascalis, and M. Pradella, "Static Analysis of Infrastructure as Code: A Survey," in Proc. IEEE 19th Int. Conf. on Software Architecture Companion (ICSA-C), 2022, pp. 218-225.
M. Guerriero, M. Garriga, D. A. Tamburri, and F. Palomba, "Adoption, Support, and Challenges of Infrastructure-as-Code: Insights from Industry," in Proc. IEEE Int. Conf. on Software Maintenance and Evolution (ICSME), 2019, pp. 580-589.
Center for Internet Security, "CIS Amazon Web Services Foundations Benchmark, v2.0," 2023.
A. Dubey et al., "The Llama 3 Herd of Models," arXiv:2407.21783, 2024.
BridgeCrew, "TerraGoat: Vulnerable Terraform Infrastructure," 2024. [Online]. Available: https://github.com/bridgecrewio/terragoat