The hardest part of the EU AI Act is no longer understanding that rules exist. It is producing credible evidence that an organization knows which AI systems it uses, what those systems are designed to do, who controls them, how they were assessed and what happens when their performance changes.
That makes EU AI Act compliance an operational discipline, not a legal document or a one-time product review. Organizations supplying or using AI in the European Union will need records that can survive scrutiny from customers, regulators, auditors and their own leadership. A broad AI policy may set direction, but it cannot answer the more practical questions: Which model was used in this hiring tool? What data reaches it? Who approved its purpose? What testing was performed? What changed after the last release?
The regulation’s risk-based structure is often described as straightforward. Its implementation is not. Legal categories must be translated into inventories, workflow decisions, technical documentation, testing programmes and clear lines of accountability across teams that do not ordinarily work from the same evidence.
A phased law, with different duties for different actors
The AI Act entered into force on 1 August 2024, but its provisions apply in stages. Rules prohibiting specified AI practices, along with AI-literacy obligations, began applying on 2 February 2025. Obligations for providers of general-purpose AI models began applying on 2 August 2025, with particular transition arrangements for models already on the market. Most provisions, including many rules for high-risk AI systems listed in Annex III and transparency obligations, apply from 2 August 2026. Requirements for certain high-risk systems connected to regulated products apply from 2 August 2027.
Timing matters, but so does role. A provider develops an AI system or has one developed and places it on the market or puts it into service under its name. A deployer uses a system under its authority. Importers and distributors have separate duties when bringing systems into the EU market or making them available. An organization can occupy more than one role—and a company that substantially modifies a system or changes its intended purpose can take on provider responsibilities.
This is why a generic statement that a business is “using third-party AI” is rarely enough. The relevant obligations depend on the system, the intended purpose, where it is used and the organization’s role in the value chain.
The AI system inventory is the starting point
The first implementation challenge is discovering what exists. A useful AI system inventory is not a spreadsheet of chatbot subscriptions. It is a living register that connects the technology to its operational context.
- The model, application, vendor and version in use
- The business use case, intended purpose and affected people
- System owners, approvers and human users
- Inputs, outputs, data sources, integrations and data flows
- Whether the tool makes, supports or materially influences decisions
- The applicable legal category and the reasoning behind it
- Testing, monitoring, incident and change-management records
Inventory work often reveals that the same underlying model is used in several products with very different consequences. A language model helping staff draft internal summaries may present a different risk profile from the same model used to rank job applicants, assess creditworthiness or support access to an essential service. The AI Act regulates systems in context; organizations should govern them that way too.
Classification must become a recorded decision
Teams must assess whether a use falls within a prohibited practice, a transparency obligation, high-risk requirements or a category with no AI Act duty specific to that use. The Act’s prohibitions are targeted rather than a blanket ban on harmful or inaccurate AI. High-risk classifications include certain safety components of regulated products and systems used in listed areas such as employment, education, access to essential private and public services, law enforcement, migration and border management, and the administration of justice and democratic processes.
Annex III classifications include qualifications and exceptions. For example, an Annex III system may not be high risk where it performs a narrow procedural task and does not materially influence decision-making, subject to conditions in the Act. Providers making that determination need to document it; deployers should not treat a vendor’s label as a substitute for their own understanding of the deployment.
The most defensible approach is to preserve the classification rationale: the intended purpose considered, facts relied on, accountable decision-maker, date of review and circumstances that would trigger reassessment. That record matters when a product is repurposed, connected to new data or introduced into a consequential workflow.
Policies are not evidence
For high-risk AI systems, the Act sets out a substantial compliance framework. Providers must establish a risk-management system, meet data and data-governance requirements where training, validation and testing datasets are used, prepare technical documentation, keep logs, provide instructions for use, enable human oversight and meet requirements relating to accuracy, robustness and cybersecurity. They must also operate a quality-management system and carry out the relevant conformity assessment before placing a system on the market or putting it into service.
In practice, an organization’s evidence trail may include:
- Risk assessments and mitigation decisions, including residual risks
- AI testing documentation covering performance, foreseeable misuse, robustness and security-relevant failure modes
- Data-governance records, including data suitability and known limitations where applicable
- Human-oversight instructions, training and escalation procedures
- System logs, incident reports, monitoring findings and corrective actions
- User-facing instructions, transparency notices and change histories
- Supplier assessments, contractual commitments and records of validation performed locally
A model card, benchmark result or supplier marketing claim can be useful input. It is not automatically sufficient technical documentation for a particular AI system or deployment. Testing must relate to the intended purpose and reasonably foreseeable conditions of use. A system that performs well in a vendor demonstration may fail when presented with an organization’s language mix, workflow constraints, user behaviour or data quality.
Third-party AI creates a handoff problem
Much of the difficult work sits between organizations. A provider may have detailed knowledge of model development but limited visibility into a customer’s local deployment. A deployer may understand the real-world decision process but lack access to the model’s training data or architecture. Procurement, legal, engineering, security, HR and operations must therefore establish a practical handoff: what the supplier can evidence, what the customer must test, who monitors performance and who acts when something goes wrong.
For deployers of high-risk systems, duties include following provider instructions, assigning human oversight to competent people, monitoring operation and, where appropriate, keeping logs under their control. Other obligations can arise in particular settings, including workplace use and public-sector deployments. The Act also operates alongside the GDPR, product-safety law, cybersecurity requirements and sector-specific rules; compliance with one does not eliminate duties under another.
General-purpose AI models add another layer. Providers of such models have their own documentation, information-sharing and policy duties, with additional obligations for models presenting systemic risk. But a downstream company still needs to evaluate its own application and use case. Supplier paperwork can inform that analysis; it cannot settle every question for the deployer.
The auditability test
Organizations preparing for European AI regulation should be able to answer a short set of questions consistently and with records:
- What AI systems and models are in use, including embedded vendor tools?
- What is each system’s intended purpose, and who is affected?
- Which AI Act category applies, and why?
- Who owns the system and approved its deployment?
- What was tested, under what conditions, and what limitations were found?
- What changed after approval, and when does that change require reassessment?
- How are users informed, humans able to intervene and incidents escalated?
- What happens if performance degrades, data shifts or a supplier changes the model?
The AI Act includes significant penalties: the highest tier can reach €35 million or 7% of worldwide annual turnover for violations involving prohibited practices, with lower tiers for other violations and for supplying incorrect information. Yet the more durable reason to build evidence now is not the prospect of a fine. It is that AI systems are increasingly embedded in decisions that organizations must explain.
EU AI Act implementation is therefore becoming a test of institutional memory. The organizations best placed to comply will not be those with the longest policy statements. They will be those that can show, system by system, how a legal obligation became a repeatable process—and how that process continues after deployment.