Source profileQuality 95/100Review permissions

brucesongs/kali-claw/skills/eu-ai-act-compliance-redteam/SKILL.md

eu-ai-act-compliance-redteam

EU AI Act (Regulation (EU) 2024/1689) compliance-focused red team testing for high-risk AI systems — Article 9 adversarial testing, Annex III classification, Annex IV technical documentation, conformity assessment, and Notified Body audit preparation. Enforceable since 2 August 2026 with fines up to €35M or 7% global turnover.

Source repository stars
65
Declared platforms
2
Static risk flags
2
Last source update
2026-08-19
Source checked
2026-08-25

Decision brief

What it does: where it fits

Supplementary Files: - payloads.md — Article 9 testing commands, Annex IV documentation generators, conformity assessment scripts, Notified Body audit prep check - test-cases.md — 5 structured test cases covering LLM red team, deepfake classification, bias testing, transparency…

Best for

  • Pre-audit readiness assessment for an enterprise deploying a high-risk LLM (e.g., HR resume screening) — generate the Article 9 evidence trail and identify gaps before a Notified Body arrives
  • Annex III high-risk classification — determine if a new AI feature (e.g., AI-driven loan approval) falls under Annex III scope and which Article 9 sub-clauses apply
  • Adversarial robustness testing (Article 15) — run documented red team suites against an in-production model, capturing inputs/outputs/metrics for the technical file

Not for

  • "Check-box" red teaming — running Garak once and calling it Article 9 compliance. Auditors reject this — testing must be per-release and per-change, with documented residual risk acceptance
  • Conflating "high-risk" with "important" — only Annex III categories count. An "important" AI for the business that isn't in Annex III is unregulated

Compatibility matrix

Platform support, with evidence labels

PlatformStatusEvidenceWhat to check
CodexNot declaredNo explicit evidencePortability before use
Claude CodeDeclaredSource recordInstall path and trigger
CursorDeclaredSource recordInstall path and trigger
Gemini CLINot declaredNo explicit evidencePortability before use
Open the compatibility checker

Installation

Inspect first. Install second.

The source command is displayed only when detected. A safe inspection prompt is always available so your agent can explain every action before execution.

Source-detected install commandSource
npx skills add https://github.com/brucesongs/kali-claw --skill "skills/eu-ai-act-compliance-redteam"
Safe inspection promptEditorial

Inspect the Agent Skill "eu-ai-act-compliance-redteam" from https://github.com/brucesongs/kali-claw/blob/a3205f5484ca8fec9fd809f3c16fe41fbc6ac87e/skills/eu-ai-act-compliance-redteam/SKILL.md at commit a3205f5484ca8fec9fd809f3c16fe41fbc6ac87e. List every install step, command, network request, credential, file read/write, external action, and rollback step. Explain whether it fits my task. Do not install or execute anything until I approve.

Workflow

What the source asks the agent to do

  1. 01

    Phase 1: Classification & Scope

    Determine whether the AI system falls under EU AI Act high-risk (Annex III) or is otherwise regulated (e.g., GPAI with systemic risk under Article 51).

    Determine whether the AI system falls under EU AI Act high-risk (Annex III) or is otherwise regulated (e.g., GPAI with systemic risk under Article 51).
  2. 02

    Phase 2: Risk Assessment & Test Plan

    Using NIST AI RMF + EU AI Act Article 9, build a risk register and corresponding test plan. Each risk maps to one or more Article 9 sub-clauses.

    Using NIST AI RMF + EU AI Act Article 9, build a risk register and corresponding test plan. Each risk maps to one or more Article 9 sub-clauses.
  3. 03

    Phase 3: Adversarial Testing Execution

    Run the test plan using Garak, Counterfit, TextAttack, Aequitas. Capture: - Input + expected output + actual output + metric (passed/failed) - Reproducibility (model commit hash, random seed, environment) - Time-stamped logs (for Article 12 / Annex IV §2(f))

    Input + expected output + actual output + metric (passed/failed)Reproducibility (model commit hash, random seed, environment)Time-stamped logs (for Article 12 / Annex IV §2(f))
  4. 04

    Phase 4: Documentation Generation

    Generate Annex IV technical file, Model Card, Datasheet, and risk register. Each artifact cross-references the test evidence captured in Phase 3.

    Generate Annex IV technical file, Model Card, Datasheet, and risk register. Each artifact cross-references the test evidence captured in Phase 3.
  5. 05

    Phase 5: Conformity Assessment & Post-Market

    Submit Annex IV file + internal control documentation for conformity assessment (internal control for most Annex III systems, Notified Body for some). Establish post-market monitoring loop (Article 72).

    Submit Annex IV file + internal control documentation for conformity assessment (internal control for most Annex III systems, Notified Body for some). Establish post-market monitoring loop (Article 72).

Permission review

Static risk signals and limitations

Writes files

medium · line 63

The documentation asks the agent to create, modify, or delete local files.

**Post-market monitoring (Article 72)** — establish the loop: production telemetry → anomaly detection → re-test → update risk register → file revision

Runs scripts

medium · line 87

The documentation asks the agent to run terminal commands or scripts.

python3 skills/eu-ai-act-compliance-redteam/scripts/classify.py \

Runs scripts

medium · line 124

The documentation asks the agent to run terminal commands or scripts.

python3 skills/eu-ai-act-compliance-redteam/scripts/annex_iv.py \

Evidence record

Why each signal appears

EvidenceSourceComputedTestedEditorial
SignalValueEvidence typeMeaning
Quality score95/100ComputedDocumentation, specificity, maintenance, and trust rules
Repository stars65SourceRepository attention, not individual Skill quality
Compatibility2 platformsSourceDeclared in the catalog source record
Usage guideautomated source guideEditorialGenerated or reviewed according to the visible evidence level

Pinned source

Provenance and original SKILL.md

Repository
brucesongs/kali-claw
Skill path
skills/eu-ai-act-compliance-redteam/SKILL.md
Commit
a3205f5484ca8fec9fd809f3c16fe41fbc6ac87e
License
MIT
Collected
2026-08-25
Default branch
main
View the original SKILL.md

Skill: EU AI Act Compliance Red Team

Supplementary Files:

  • payloads.md — Article 9 testing commands, Annex IV documentation generators, conformity assessment scripts, Notified Body audit prep check
  • test-cases.md — 5 structured test cases covering LLM red team, deepfake classification, bias testing, transparency docs, and robustness adversarial
  • guides/eu-ai-act-article-9-deep-dive.md — Article 9 clause-by-clause analysis with test mapping

Summary

EU AI Act compliance-focused red team skill domain. Validating that high-risk AI systems satisfy Article 9 (data governance, adversarial robustness, logging, transparency, human oversight) and Annex IV (technical documentation) requirements, with conformity assessment evidence and Notified Body audit preparation.

Domain: ai-compliance | Regulation: EU Regulation 2024/1689 | Enforcement: since 2026-08-02

Description

The EU AI Act (Regulation (EU) 2024/1689) is the world's first comprehensive legal framework for artificial intelligence. After a phased implementation beginning February 2025, the high-risk AI system obligations under Articles 8–15 entered enforceable status on 2 August 2026. National competent authorities (the AI Office plus each Member State's regulator) may now initiate audits, demand technical documentation, and impose penalties.

The maximum penalty: €35 million or 7% of global annual turnover, whichever is higher. For mid-cap enterprises the threshold drops to €15M / 3%, but for systemic-risk general-purpose AI models the maximum still applies.

Article 9 specifically mandates that providers of high-risk AI systems conduct adversarial testing (red teaming) to identify and mitigate known and reasonably foreseeable vulnerabilities. This is distinct from technical AI red teaming (covered in ai-safety-redteam-advanced) — Article 9 testing must be documented in Annex IV technical files, with risk treatment decisions, residual risk acceptance, and retest schedules.

This skill covers the compliance perspective: how to map a high-risk AI system to Annex III categories, run red-team tests that satisfy Article 9 acceptance criteria, generate Annex IV technical documentation artifacts, prepare for conformity assessment, and survive a Notified Body audit.

Skill Identity

AspectValue
TypeCompliance testing + documentation generation
Distinguishing featureRegulatory lens applied to technical red team work
Adjacent skillsai-safety-redteam-advanced (technical), ai-security (model/app security), llm-red-team (LLM-specific)
Distinct from adjacentThis skill produces compliance evidence; adjacent skills produce vulnerability findings

Why this skill exists

Three converging trends made this skill necessary in 2026:

  1. Enforcement began 2026-08-02. Annex III high-risk systems (biometrics, critical infrastructure, education, employment, essential services, law enforcement, migration, justice, democratic processes) must comply. Companies that did not prepare Article 9 evidence in 2025-2026 are now exposed to fines.
  2. Annex IV technical documentation is the audit deliverable. Most technical red team tools (Garak, Counterfit, TextAttack) produce findings; Annex IV requires evidence-attached, traceable, retested-on-schedule findings. There is a documentation gap.
  3. Notified Body audits in late 2026 will be the first wave. Companies facing audit need a structured way to organize the artifacts they have (or generate what's missing).

Differentiation from ai-safety-redteam-advanced

Dimensionai-safety-redteam-advancedThis skill
GoalFind technical vulnerabilitiesProduce regulatory compliance evidence
OutputVulnerability report, CVSS scores, PoCsAnnex IV technical file, risk register, audit trail
AudienceSecurity engineers, CISONotified Body auditor, AI Office regulator, compliance team
ToolsOWASP LLM Top 10, jailbreaks, prompt injectionArticle 9 test matrix, Annex IV template, conformity checklists
Test cadenceContinuous / on-demandPer-release + per-change + per-retest-window

Both skills use overlapping tools (Garak, Counterfit, TextAttack) but produce different artifacts for different audiences.

Use Cases

  1. Pre-audit readiness assessment for an enterprise deploying a high-risk LLM (e.g., HR resume screening) — generate the Article 9 evidence trail and identify gaps before a Notified Body arrives
  2. Annex III high-risk classification — determine if a new AI feature (e.g., AI-driven loan approval) falls under Annex III scope and which Article 9 sub-clauses apply
  3. Adversarial robustness testing (Article 15) — run documented red team suites against an in-production model, capturing inputs/outputs/metrics for the technical file
  4. Data governance audit (Article 10) — verify training data provenance, bias mitigation, andDatasheet generation completeness
  5. Transparency documentation (Article 13) — generate user-facing notices, instructions for use, and confidence-score disclosures for deployers
  6. Human oversight effectiveness (Article 14) — test that human reviewers can override AI outputs and detect when the AI is wrong; log override rates and decision quality
  7. Logging completeness (Article 12 / Annex IV §2(f)) — verify automatic event logging captures training runs, validation tests, post-market monitoring events
  8. Conformity assessment preparation — assemble Annex IV technical documentation (system architecture, model card, training data summary, test results, risk register)
  9. Post-market monitoring (Article 72) — establish the loop: production telemetry → anomaly detection → re-test → update risk register → file revision
  10. Cross-border deployment check — verify that a model trained and approved in one Member State is acceptable for deployment in another (mutual recognition under Article 60)

Core Tools

ToolCategoryPurposeLicense
Garak (NVIDIA)LLM vulnerability scannerProbe LLMs for OWASP LLM Top 10 + EU AI Act Article 15 robustness; CSV/JSON output for Annex IVApache 2.0
Microsoft CounterfitAI red-team orchestrationRun adversarial suites against ML models (vision, NLP, tabular); task-level evidence collectionMIT
TextAttackNLP adversarialGenerate and run adversarial text perturbations for robustness testingMIT
Audit-MLModel decision auditLog feature attribution + outcome explanation; supports Article 13 transparencyApache 2.0
Aequitas / FairlearnBias auditingQuantify demographic disparity in model outcomes; supports Article 10 data governanceMIT / MIT
Model Card Toolkit (Google)Documentation generationGenerate Model Cards (Mitchell et al. 2019) that map to Annex IV §2(c)Apache 2.0
Datasheets for Datasets templatesDocumentation generationFill Gebru et al. 2018 datasheet template → Annex IV §2(a)CC-BY
NIST AI RMF (framework)Risk management alignmentMap EU AI Act evidence to NIST AI RMF Govern/Map/Measure/Manage for cross-jurisdictionPublic domain

Methodology

Phase 1: Classification & Scope

Determine whether the AI system falls under EU AI Act high-risk (Annex III) or is otherwise regulated (e.g., GPAI with systemic risk under Article 51).

# Decision-tree prompt
python3 skills/eu-ai-act-compliance-redteam/scripts/classify.py \
  --system-description "LLM screening resumes for senior engineer role" \
  --output scope.json
# Outputs: { "annex_iii_clause": "IV(a)", "high_risk": true, "obligations": ["Art.9","Art.10","Art.13","Art.14","Art.15"] }

Phase 2: Risk Assessment & Test Plan

Using NIST AI RMF + EU AI Act Article 9, build a risk register and corresponding test plan. Each risk maps to one or more Article 9 sub-clauses.

# Risk register entry
risk = {
    "id": "R-AI-2026-001",
    "description": "Model fails on non-English EU official language resumes",
    "article_9_clause": "Art.9(2)(b) — error-free performance across population groups",
    "severity": "high",
    "test_suite": "multilingual_resume_eval",
    "mitigation": "fine-tune on Spanish/Polish/German HR data; add pre-deployment language coverage test",
    "residual_risk_acceptance": "low — CEO signoff required",
    "retest_cadence": "quarterly"
}

Phase 3: Adversarial Testing Execution

Run the test plan using Garak, Counterfit, TextAttack, Aequitas. Capture:

  • Input + expected output + actual output + metric (passed/failed)
  • Reproducibility (model commit hash, random seed, environment)
  • Time-stamped logs (for Article 12 / Annex IV §2(f))

Phase 4: Documentation Generation

Generate Annex IV technical file, Model Card, Datasheet, and risk register. Each artifact cross-references the test evidence captured in Phase 3.

# Generate Annex IV technical file (Markdown + JSON)
python3 skills/eu-ai-act-compliance-redteam/scripts/annex_iv.py \
  --risk-register risks.json \
  --test-evidence evidence/ \
  --model-card model_card.json \
  --datasheet data/datasheet.yaml \
  --output annex_iv_technical_file.md

Phase 5: Conformity Assessment & Post-Market

Submit Annex IV file + internal control documentation for conformity assessment (internal control for most Annex III systems, Notified Body for some). Establish post-market monitoring loop (Article 72).

Defense Perspective

Defense LayerControlKey Points
Classification (Phase 1)Document Annex III reasoning in scope.json; require legal review for borderline casesMisclassification is the #1 audit failure mode — every Notified Body will challenge the classification first
Risk assessment (Phase 2)Use both top-down (scenario-based) and bottom-up (component-based) risk ID; cross-walk to NIST AI RMFArticle 9 doesn't prescribe a methodology; choosing a recognized framework (NIST AI RMF, ISO/IEC 23894) reduces audit pushback
Test execution (Phase 3)Reproducibility: pin model commit, random seeds, dependency versions; store raw outputs alongside processed resultsAuditors request re-runs; non-reproducible test = invalid evidence
Documentation (Phase 4)Annex IV file should be machine-readable (JSON) AND human-readable (Markdown); generate both from same sourceAuditors use both formats depending on whether they are technical or legal
Conformity (Phase 5)Internal-control conformity (Annex III default) needs 4 documents: technical file, quality management system, EU declaration of conformity, CE mark registration; Notified Body conformity (Article 6(3)) needs 7 documentsDon't conflate the two conformity paths — the documentation burden differs
Post-market (Art. 72)Logging requirements under Article 12 + Annex IV §2(f) require automatic capture of training runs, post-deployment events, and serious incidents; serious incidents must be reported to the AI Office within 15 days (Article 73)Logging must be designed, not bolted on; retrofit is painful
Penalty mitigationMaintain evidence that provider exercised "due diligence" — Article 9 evidence + corrective action on past findings = strongest mitigation argumentThe AI Office has discretion; documented good faith is the strongest defense

Detection Methods

Detection of non-compliance with the EU AI Act — the regulator's perspective.

Regulator Detection (What auditors look for)

  • Absence of risk register — first thing an auditor asks for; missing = immediate red flag
  • Test evidence older than 12 months — Article 9 implies currency; stale tests = non-compliance
  • Annex IV file without traceability — claims without reproducibility (no commit hash, no seed) = invalid
  • Missing serious-incident reports — Article 73 requires reporting within 15 days; auditor checks registry
  • Inadequate logging — Article 12 + Annex IV §2(f) define specific event types; missing any = finding
  • No post-market monitoring plan — Article 72 requires a documented process; absence = non-compliance

Provider Self-Detection

# Compliance self-check
python3 skills/eu-ai-act-compliance-redteam/scripts/self_check.py --model my_model

# Sigma rule for missing Annex IV artifacts (run in CI)
# Detects: PR that trains a model without Annex IV file update
title: Model trained without Annex IV update
logsource:
  product: ci
detection:
  selection:
    event_type: model_train
    annex_iv_changed: false
  condition: selection

Defense Evasion Techniques

How a non-compliant provider might evade detection (for awareness, not endorsement):

  1. Cosmetic Annex IV — generate a technically-complete Annex IV file that makes claims unsupported by actual test evidence; auditors catch this by requesting raw evidence
  2. Outdated test reports — present last year's tests as current; mitigated by date checks + re-test requests
  3. "Continuous improvement" misdirection — claim the model is continuously improved so individual version tests don't apply; AI Office guidance (2026-Q1) explicitly rejected this argument
  4. GPAI classification avoidance — claim a model with 10^25 FLOPs is below the systemic-risk threshold; auditors will demand training records
  5. Borderline Annex III misclassification — claim a hiring LLM is "assistance" not "filtering" (Annex III IV(a)); regulators have published 12 precedent rulings in 2026-Q2 narrowing this
  6. Transparency washing — user-facing notices buried in 50-page ToS; Article 50 requires "clear and conspicuous" disclosure

Practical Steps

Detailed payloads in payloads.md, complete test checklist in test-cases.md.

Step 1: Classify the AI system

Run classify.py (Phase 1) to determine Annex III scope. If borderline, escalate to legal review.

Step 2: Build the risk register

Use the risk_register_template.json; populate top-down (scenarios) and bottom-up (components).

Step 3: Execute the test plan

For LLMs: Garak + TextAttack. For vision: Counterfit. For tabular: Aequitas + Audit-ML.

Step 4: Generate Annex IV

Run annex_iv.py (Phase 4); review the generated Markdown; sign off with risk owner.

Step 5: Submit for conformity assessment

Internal control (default): self-attest with EU declaration of conformity + CE mark. Notified Body (Article 6(3)): engage auditor.

Step 6: Operate post-market monitoring

Establish logging per Article 12; configure serious-incident reporting pipeline (15-day deadline).

Common Pitfalls

  • "Check-box" red teaming — running Garak once and calling it Article 9 compliance. Auditors reject this — testing must be per-release and per-change, with documented residual risk acceptance
  • Conflating "high-risk" with "important" — only Annex III categories count. An "important" AI for the business that isn't in Annex III is unregulated
  • GPAI threshold confusion — the 10^25 FLOPs systemic-risk threshold (Article 51) is different from high-risk (Annex III). A model can be both, neither, or one
  • Neglecting logging design — bolting on Article 12 logging after deployment is 10x more expensive than designing it in
  • Missing 15-day serious-incident window — Article 73 clock starts when provider becomes aware, not when incident is confirmed. Establish a triage SLA
  • Cross-border mutual recognition assumed — Article 60 mutual recognition exists but Member State regulators have challenged it in practice for politically sensitive deployments (e.g., law enforcement)
  • Transparency doc as afterthought — Article 13 transparency is a deployer obligation, not just a provider one. Deployers (companies using the AI) are separately liable

Cross-Reference to Related Skills

  • ai-safety-redteam-advanced — technical AI red team (OWASP LLM Top 10)
  • ai-agent-security — agent framework attacks
  • llm-red-team — LLM-specific attacks
  • ai-fuzzing — ML fuzzing tools
  • ci-cd-supply-chain-attack — supply chain compliance overlap
  • secret-management-attack — credential handling for AI systems

Hacker Laws Alignment

  • Law 4 (Verify Everything): Article 9 evidence must be reproducible; trust nothing that isn't traceable to a commit hash
  • Law 7 (Documentation is Part of the System): Annex IV is part of the AI system, not an add-on

References

Attribution

This skill codifies EU AI Act compliance red team practice as of 2026-08. Regulation continues to evolve through EU implementing acts (expected late 2026) and Member State national law. Always verify against the latest EUR-Lex publication before relying on this skill for legal compliance.

Frequently asked questions

What to verify before installation and use

What does the eu-ai-act-compliance-redteam source document cover?

Supplementary Files: - payloads.md — Article 9 testing commands, Annex IV documentation generators, conformity assessment scripts, Notified Body audit prep check - test-cases.md — 5 structured test cases covering LLM red team, deepfake classification, bias testing, transparency…

How do I install eu-ai-act-compliance-redteam?

The source record exposes this install command: npx skills add https://github.com/brucesongs/kali-claw --skill "skills/eu-ai-act-compliance-redteam". Inspect the command and pinned source before running it.

Which Agent platforms does the source record declare?

The pinned source record declares support for: claude code, cursor.

Which permission-related actions were detected?

Static rules flagged write-files, exec-script in the source; the page lists the matching lines and excerpts.

Alternatives

Compare before choosing

Computed 967

aomi-labs/skills

aomi-build

Scaffold new Aomi apps and plugins from API docs, OpenAPI/Swagger specs, or SDK references. aomi-build generates production-ready Rust SDK crates (lib.rs, client.rs, tool.rs) with tool schemas, preambles, host-interop flows, and validation — turning a vendor's API surface into AI-agent-callable tools. It covers the current `aomi-build` OpenAPI pipeline (`gen-specs` → `gen-client` → `gen-tool` → curate → compile/test) as well as greenfield apps. Use when the user wants to scaffold a new Aomi app

Computed 95203

PramodDutta/qaskills

API Test Suite Generator

Automatically generate comprehensive API test suites from OpenAPI specifications covering CRUD operations, error handling, authentication, pagination, and edge cases

Computed 94203

PramodDutta/qaskills

Code Review Excellence

Master code review best practices with constructive feedback patterns, quality assurance standards, review checklists, security considerations, and collaborative improvement techniques for high-quality software delivery.

Computed 9420

upex-galaxy/agentic-qa-boilerplate

test-automation

Plan, write, and review automated tests following KATA (Komponent Action Test Architecture) on Playwright + TypeScript, or explain existing automated tests in a sealed read-only mode. Use when writing E2E or API/integration tests, creating Page or Api components, designing ATCs, parameterizing test data, registering fixtures, reviewing test code for KATA compliance, or requesting break-down-tests / a plain-English test breakdown. The explain mode reads source and reports assertions without enter