Article 15
EU AI Act Article 15 — Accuracy, Robustness and Cybersecurity
Article 15 of Regulation (EU) 2024/1689 requires high-risk AI systems to achieve appropriate accuracy, robustness and cybersecurity throughout their lifecycle. Disaggregated metrics, adversarial testing, and what to document.
Source: Regulation (EU) 2024/1689 on EUR-Lex · Last published 2026-04-28 · Draft pending human review
What Article 15 actually requires
Article 15 of Regulation (EU) 2024/1689 sets three intertwined requirements for high-risk AI systems: accuracy, robustness and cybersecurity. The system must be designed and developed to achieve appropriate levels of all three throughout its lifecycle, and these levels (with the relevant accuracy metrics) must be declared in the Article 13 instructions for use.
Reg text — Article 15(1): "High-risk AI systems shall be designed and developed in such a way that they achieve an appropriate level of accuracy, robustness and cybersecurity, and that they perform consistently in those respects throughout their lifecycle."
The three pillars
- Accuracy. Article 15(3) requires accuracy levels and metrics to be declared. Annex IV(2)(b) requires metrics disaggregated across the persons or groups in relation to whom the system is intended to be used. Single-number accuracy is insufficient.
- Robustness. Article 15(4) requires resilience against errors, faults or inconsistencies that may occur in the system or the environment, including with regard to feedback loops where the system continues to learn after being placed on the market.
- Cybersecurity. Article 15(5) requires resilience against attempts by unauthorised third parties to alter behaviour or performance — data poisoning, model poisoning, model evasion, confidentiality attacks, prompt injection where applicable.
The Commission, in cooperation with relevant stakeholders and organisations such as ENISA, may develop benchmarks and measurement methodologies (Article 15(2)).
Who is covered
Providers design these properties into the system. Deployers exercise them through input-data quality (Article 26(4)) and operational monitoring (Article 26(5)).
What to do
- Run a validation suite that produces disaggregated metrics by demographic group on the validation and test sets.
- Run adversarial test campaigns before each release, covering at minimum data poisoning, model evasion, model extraction, membership inference, and (for LLM-based features) prompt injection. Document results in Annex IV §2(g).
- Implement drift monitoring in production tied to the Article 9 post-market loop.
- Document feedback-loop guards: continuous learning systems must be analysed for amplification of bias or drift.
- Cite the harmonised standards you align with — ISO/IEC 24029-2:2023 (robustness), ISO/IEC 27001 + sector-specific cybersecurity standards.
Inline crosswalk
- ISO/IEC 42001:2023 Annex A.6.2.4 — Verification and validation.
- ISO/IEC 42001:2023 Annex A.6.2.5 — Security of AI systems.
- NIST AI RMF MEASURE 2.5 — System is demonstrated to be valid and reliable.
- NIST AI RMF MEASURE 2.7 — System security and resilience are evaluated and documented.
Common mistakes
- Single-number accuracy. Article 15(3) + Annex IV(2)(b) require disaggregated metrics.
- Cybersecurity testing limited to web-app penetration testing. Article 15(5) wants ML-specific attack vectors.
- No drift monitoring after launch.
- Continuous-learning loops without bias amplification analysis.
Penalties
Article 99(4) — up to €15 million or 3% of worldwide annual turnover.
Disclaimer. Reference; not legal advice. Verify with counsel. Reg text from Regulation (EU) 2024/1689.
Reference checklist
From the Governancer 30-item EU AI Act checklist. Each item joins to the ISO 42001 + NIST AI RMF crosswalk table below.
Article 15 · Starter tier · medium
Define accuracy, robustness, and cybersecurity measures
Benchmarks, adversarial testing, incident detection, lifecycle monitoring. Documented in the technical file.
Annex IV · Pro tier · high
Adversarial-testing results against common attack vectors
Annex IV(2)(g) requires documentation of cybersecurity measures. Cover data poisoning, model extraction, evasion, and prompt injection with real test results.
Annex IV · Pro tier · critical
Disaggregated performance metrics by demographic group
Article 15(3) + Annex IV(2)(b) require accuracy metrics reported across the groups on which the system is intended to be used — not just overall averages.
ISO 42001 + NIST AI RMF crosswalk
Pulled live from the Governancer crosswalk module. Mapping reference; not a substitute for ISO 42001 certification audit or NIST AI RMF self-attestation.
ISO/IEC 42001:2023
| Checklist item | ISO 42001 control | Rationale |
|---|---|---|
art15-accuracy | ISO/IEC 42001:2023 Annex A.6.2.4 — Verification and validation | Article 15 accuracy/robustness benchmarks satisfy the verification-and-validation control objective in the lifecycle annex. |
annexiv-cybersecurity | ISO/IEC 42001:2023 Annex A.6.2.5 — Security of AI systems | Adversarial-testing results against poisoning, extraction, evasion and prompt injection evidence the AI-security control of Annex A.6.2.5. |
annexiv-performance-groups | ISO/IEC 42001:2023 Annex A.6.2.4 — Verification and validation | Disaggregated metrics by demographic group are the validation-across-intended-population evidence required by Annex A.6.2.4. |
NIST AI RMF 1.0
| Checklist item | NIST AI RMF subcategory | Rationale |
|---|---|---|
art15-accuracy | NIST AI RMF MEASURE 2.5 — The AI system to be deployed is demonstrated to be valid and reliable | Article 15 accuracy benchmarks and robustness evidence demonstrate validity and reliability per MEASURE 2.5. |
art15-accuracy | NIST AI RMF MEASURE 2.7 — AI system security and resilience are evaluated and documented | Article 15 cybersecurity measures and adversarial resilience are the security/resilience evaluation in MEASURE 2.7. |
annexiv-cybersecurity | NIST AI RMF MEASURE 2.7 — AI system security and resilience are evaluated and documented | Adversarial-testing results across attack vectors are the documented security/resilience evaluation MEASURE 2.7 calls for. |
annexiv-performance-groups | NIST AI RMF MEASURE 2.11 — Fairness and bias — as identified in the MAP function — are evaluated and results are documented | Disaggregated performance metrics by demographic group are the fairness measurement MEASURE 2.11 demands. |
Related
Article 9
EU AI Act Article 9 — Risk Management System Requirements
Article 10
EU AI Act Article 10 — Data and Data Governance
Article 11
EU AI Act Article 11 — Technical Documentation Requirements
Article 12
EU AI Act Article 12 — Record-Keeping and Automatic Logging
Annex IV §2(b) + §3
Annex IV — Accuracy Documentation (EU AI Act)
Annex IV §2(d)
Annex IV §2(d) — Data Documentation (EU AI Act)
Pro feature
Generate Article 11 with AI
LLM-assisted draft of all eight Annex IV sections, pre-filled from your system intake. 5 drafts/month on Pro.
Pro template
Download FRIA template
15-page Article 27 FRIA template (.docx) with the six elements pre-structured and a worked example.
Get the 30-item EU AI Act compliance checklist
Free PDF. No spam. Maps every Article and Annex IV section we ship to a ready-to-action checklist row.
Reference; not legal advice. Verify with qualified counsel before relying on it for compliance decisions. Reg text quoted from the Official Journal version of Regulation (EU) 2024/1689. Published by Agonist Development AB.