Annex IV §2(b) + §3
Annex IV — Accuracy Documentation (EU AI Act)
How to document accuracy, robustness and cybersecurity in the EU AI Act technical file. Disaggregated metrics, adversarial testing, and the link to Article 15.
Source: Regulation (EU) 2024/1689 on EUR-Lex · Last published 2026-04-28 · Draft pending human review
What the accuracy paragraphs actually require
The technical-file rendering of Article 15 sits across two Annex IV provisions:
- §2(b) — Description of the methods, design specifications and validation, including the algorithmic logic, the expected accuracy in relation to its intended purpose, the accuracy metrics disaggregated for relevant groups of persons.
- §3 — Detailed information about the monitoring, functioning and control of the AI system, in particular the metrics used to measure accuracy, robustness and cybersecurity, and any foreseeable unintended outcomes and sources of risk.
Article 15(3) is the load-bearing requirement: accuracy levels and metrics declared in the instructions for use, with disaggregation across the persons or groups in relation to whom the system is intended to be used.
What to include
- Validation methodology. Train/val/test split design, stratification, statistical significance.
- Aggregate accuracy with confidence intervals.
- Disaggregated accuracy by demographic / contextual sub-group. Required by Article 15(3) + Annex IV §2(b).
- Robustness testing — input perturbation, distribution shift, edge-case probing, ISO/IEC 24029-2:2023 alignment where applicable.
- Cybersecurity testing — adversarial campaigns covering data poisoning, model extraction, model evasion, membership inference, and (for LLM-based features) prompt injection. Required by Article 15(5).
- Foreseeable unintended outcomes with mitigation status. Tied to the Article 9 register.
- Drift-monitoring plan for in-deployment validation. Tied to §8 post-market monitoring.
A worked metric table
| Sub-group | Precision | Recall | False-positive rate | n |
|---|---|---|---|---|
| Overall | 0.91 | 0.87 | 0.06 | 41,200 |
| Sex = F | 0.89 | 0.85 | 0.07 | 17,800 |
| Sex = M | 0.92 | 0.88 | 0.06 | 23,300 |
| Age 18–25 | 0.86 | 0.82 | 0.09 | 6,400 |
| Age 26–55 | 0.93 | 0.89 | 0.05 | 27,200 |
| Age 56+ | 0.88 | 0.84 | 0.08 | 7,600 |
Plus a commentary paragraph: "The 7pp gap between the highest- and lowest-performing age sub-group falls outside the 5pp parity target documented in the Article 9 register row R-LEG-014. Mitigation: re-weighting in the loss function, deployed in v3.1.0; re-validation in v3.1.1 closes the gap to 3.4pp. Residual risk Low; flagged in §4 for quarterly review."
Inline crosswalk
- ISO/IEC 42001:2023 Annex A.6.2.4 — Verification and validation.
- ISO/IEC 42001:2023 Annex A.6.2.5 — Security of AI systems.
- NIST AI RMF MEASURE 2.5 — System demonstrated to be valid and reliable.
- NIST AI RMF MEASURE 2.7 — Security and resilience evaluated.
- NIST AI RMF MEASURE 2.11 — Fairness and bias evaluated and documented.
Common mistakes
- Single-number accuracy, no disaggregation. Article 15(3) requires sub-group metrics.
- Cybersecurity testing limited to web-app pen-testing. ML-specific attacks are required.
- No foreseeable-unintended-outcome paragraph in §3.
- No drift-monitoring plan tied to post-market monitoring.
Disclaimer. Reference; not legal advice. Verify with counsel. Reg text from Regulation (EU) 2024/1689.
Reference checklist
From the Governancer 30-item EU AI Act checklist. Each item joins to the ISO 42001 + NIST AI RMF crosswalk table below.
Article 15 · Starter tier · medium
Define accuracy, robustness, and cybersecurity measures
Benchmarks, adversarial testing, incident detection, lifecycle monitoring. Documented in the technical file.
Annex IV · Pro tier · high
Adversarial-testing results against common attack vectors
Annex IV(2)(g) requires documentation of cybersecurity measures. Cover data poisoning, model extraction, evasion, and prompt injection with real test results.
Annex IV · Pro tier · critical
Disaggregated performance metrics by demographic group
Article 15(3) + Annex IV(2)(b) require accuracy metrics reported across the groups on which the system is intended to be used — not just overall averages.
ISO 42001 + NIST AI RMF crosswalk
Pulled live from the Governancer crosswalk module. Mapping reference; not a substitute for ISO 42001 certification audit or NIST AI RMF self-attestation.
ISO/IEC 42001:2023
| Checklist item | ISO 42001 control | Rationale |
|---|---|---|
art15-accuracy | ISO/IEC 42001:2023 Annex A.6.2.4 — Verification and validation | Article 15 accuracy/robustness benchmarks satisfy the verification-and-validation control objective in the lifecycle annex. |
annexiv-cybersecurity | ISO/IEC 42001:2023 Annex A.6.2.5 — Security of AI systems | Adversarial-testing results against poisoning, extraction, evasion and prompt injection evidence the AI-security control of Annex A.6.2.5. |
annexiv-performance-groups | ISO/IEC 42001:2023 Annex A.6.2.4 — Verification and validation | Disaggregated metrics by demographic group are the validation-across-intended-population evidence required by Annex A.6.2.4. |
NIST AI RMF 1.0
| Checklist item | NIST AI RMF subcategory | Rationale |
|---|---|---|
art15-accuracy | NIST AI RMF MEASURE 2.5 — The AI system to be deployed is demonstrated to be valid and reliable | Article 15 accuracy benchmarks and robustness evidence demonstrate validity and reliability per MEASURE 2.5. |
art15-accuracy | NIST AI RMF MEASURE 2.7 — AI system security and resilience are evaluated and documented | Article 15 cybersecurity measures and adversarial resilience are the security/resilience evaluation in MEASURE 2.7. |
annexiv-cybersecurity | NIST AI RMF MEASURE 2.7 — AI system security and resilience are evaluated and documented | Adversarial-testing results across attack vectors are the documented security/resilience evaluation MEASURE 2.7 calls for. |
annexiv-performance-groups | NIST AI RMF MEASURE 2.11 — Fairness and bias — as identified in the MAP function — are evaluated and results are documented | Disaggregated performance metrics by demographic group are the fairness measurement MEASURE 2.11 demands. |
Related
Article 15
EU AI Act Article 15 — Accuracy, Robustness and Cybersecurity
Article 10
EU AI Act Article 10 — Data and Data Governance
Annex IV §2(d)
Annex IV §2(d) — Data Documentation (EU AI Act)
Annex IV §1
Annex IV §1 — General Description of the AI System (EU AI Act)
Article 11
EU AI Act Article 11 — Technical Documentation Requirements
Pro feature
Generate Article 11 with AI
LLM-assisted draft of all eight Annex IV sections, pre-filled from your system intake. 5 drafts/month on Pro.
Pro template
Download FRIA template
15-page Article 27 FRIA template (.docx) with the six elements pre-structured and a worked example.
Get the 30-item EU AI Act compliance checklist
Free PDF. No spam. Maps every Article and Annex IV section we ship to a ready-to-action checklist row.
Reference; not legal advice. Verify with qualified counsel before relying on it for compliance decisions. Reg text quoted from the Official Journal version of Regulation (EU) 2024/1689. Published by Agonist Development AB.