Annex IV §2(b) + §3

Annex IV — Accuracy Documentation (EU AI Act)

How to document accuracy, robustness and cybersecurity in the EU AI Act technical file. Disaggregated metrics, adversarial testing, and the link to Article 15.

Source: Regulation (EU) 2024/1689 on EUR-Lex · Last published 2026-04-28 · Draft pending human review

What the accuracy paragraphs actually require

The technical-file rendering of Article 15 sits across two Annex IV provisions:

  • §2(b) — Description of the methods, design specifications and validation, including the algorithmic logic, the expected accuracy in relation to its intended purpose, the accuracy metrics disaggregated for relevant groups of persons.
  • §3 — Detailed information about the monitoring, functioning and control of the AI system, in particular the metrics used to measure accuracy, robustness and cybersecurity, and any foreseeable unintended outcomes and sources of risk.

Article 15(3) is the load-bearing requirement: accuracy levels and metrics declared in the instructions for use, with disaggregation across the persons or groups in relation to whom the system is intended to be used.

What to include

  • Validation methodology. Train/val/test split design, stratification, statistical significance.
  • Aggregate accuracy with confidence intervals.
  • Disaggregated accuracy by demographic / contextual sub-group. Required by Article 15(3) + Annex IV §2(b).
  • Robustness testing — input perturbation, distribution shift, edge-case probing, ISO/IEC 24029-2:2023 alignment where applicable.
  • Cybersecurity testing — adversarial campaigns covering data poisoning, model extraction, model evasion, membership inference, and (for LLM-based features) prompt injection. Required by Article 15(5).
  • Foreseeable unintended outcomes with mitigation status. Tied to the Article 9 register.
  • Drift-monitoring plan for in-deployment validation. Tied to §8 post-market monitoring.

A worked metric table

Sub-groupPrecisionRecallFalse-positive raten
Overall0.910.870.0641,200
Sex = F0.890.850.0717,800
Sex = M0.920.880.0623,300
Age 18–250.860.820.096,400
Age 26–550.930.890.0527,200
Age 56+0.880.840.087,600

Plus a commentary paragraph: "The 7pp gap between the highest- and lowest-performing age sub-group falls outside the 5pp parity target documented in the Article 9 register row R-LEG-014. Mitigation: re-weighting in the loss function, deployed in v3.1.0; re-validation in v3.1.1 closes the gap to 3.4pp. Residual risk Low; flagged in §4 for quarterly review."

Inline crosswalk

  • ISO/IEC 42001:2023 Annex A.6.2.4 — Verification and validation.
  • ISO/IEC 42001:2023 Annex A.6.2.5 — Security of AI systems.
  • NIST AI RMF MEASURE 2.5 — System demonstrated to be valid and reliable.
  • NIST AI RMF MEASURE 2.7 — Security and resilience evaluated.
  • NIST AI RMF MEASURE 2.11 — Fairness and bias evaluated and documented.

Common mistakes

  • Single-number accuracy, no disaggregation. Article 15(3) requires sub-group metrics.
  • Cybersecurity testing limited to web-app pen-testing. ML-specific attacks are required.
  • No foreseeable-unintended-outcome paragraph in §3.
  • No drift-monitoring plan tied to post-market monitoring.

Disclaimer. Reference; not legal advice. Verify with counsel. Reg text from Regulation (EU) 2024/1689.

Reference checklist

From the Governancer 30-item EU AI Act checklist. Each item joins to the ISO 42001 + NIST AI RMF crosswalk table below.

  • Article 15 · Starter tier · medium

    Define accuracy, robustness, and cybersecurity measures

    Benchmarks, adversarial testing, incident detection, lifecycle monitoring. Documented in the technical file.

  • Annex IV · Pro tier · high

    Adversarial-testing results against common attack vectors

    Annex IV(2)(g) requires documentation of cybersecurity measures. Cover data poisoning, model extraction, evasion, and prompt injection with real test results.

  • Annex IV · Pro tier · critical

    Disaggregated performance metrics by demographic group

    Article 15(3) + Annex IV(2)(b) require accuracy metrics reported across the groups on which the system is intended to be used — not just overall averages.

ISO 42001 + NIST AI RMF crosswalk

Pulled live from the Governancer crosswalk module. Mapping reference; not a substitute for ISO 42001 certification audit or NIST AI RMF self-attestation.

ISO/IEC 42001:2023

Checklist itemISO 42001 controlRationale
art15-accuracyISO/IEC 42001:2023 Annex A.6.2.4 — Verification and validationArticle 15 accuracy/robustness benchmarks satisfy the verification-and-validation control objective in the lifecycle annex.
annexiv-cybersecurityISO/IEC 42001:2023 Annex A.6.2.5 — Security of AI systemsAdversarial-testing results against poisoning, extraction, evasion and prompt injection evidence the AI-security control of Annex A.6.2.5.
annexiv-performance-groupsISO/IEC 42001:2023 Annex A.6.2.4 — Verification and validationDisaggregated metrics by demographic group are the validation-across-intended-population evidence required by Annex A.6.2.4.

NIST AI RMF 1.0

Checklist itemNIST AI RMF subcategoryRationale
art15-accuracyNIST AI RMF MEASURE 2.5 — The AI system to be deployed is demonstrated to be valid and reliableArticle 15 accuracy benchmarks and robustness evidence demonstrate validity and reliability per MEASURE 2.5.
art15-accuracyNIST AI RMF MEASURE 2.7 — AI system security and resilience are evaluated and documentedArticle 15 cybersecurity measures and adversarial resilience are the security/resilience evaluation in MEASURE 2.7.
annexiv-cybersecurityNIST AI RMF MEASURE 2.7 — AI system security and resilience are evaluated and documentedAdversarial-testing results across attack vectors are the documented security/resilience evaluation MEASURE 2.7 calls for.
annexiv-performance-groupsNIST AI RMF MEASURE 2.11 — Fairness and bias — as identified in the MAP function — are evaluated and results are documentedDisaggregated performance metrics by demographic group are the fairness measurement MEASURE 2.11 demands.

Pro feature

Generate Article 11 with AI

LLM-assisted draft of all eight Annex IV sections, pre-filled from your system intake. 5 drafts/month on Pro.

Pro template

Download FRIA template

15-page Article 27 FRIA template (.docx) with the six elements pre-structured and a worked example.

Get the 30-item EU AI Act compliance checklist

Free PDF. No spam. Maps every Article and Annex IV section we ship to a ready-to-action checklist row.


Reference; not legal advice. Verify with qualified counsel before relying on it for compliance decisions. Reg text quoted from the Official Journal version of Regulation (EU) 2024/1689. Published by Agonist Development AB.