The risk of an AI system does not depend only on the model: it emerges from its execution environment

In debates around the AI Act, a recurring confusion appears: the risk of an AI system is sometimes discussed as if it were mainly contained inside the model itself.

In practice, however, two systems using exactly the same model can present radically different levels of risk depending on their context of use, the data they can access, the tools they are connected to, the level of autonomy granted, and the possible consequences for people.

The model matters.
But it is not enough to determine the real operational risk.

The AI Act already recognises this logic through a risk-based approach and by classifying certain systems as “high-risk” when they are used in sensitive domains. Annex III notably covers biometrics, critical infrastructure, education, employment, essential services, law enforcement, migration, justice and democratic processes.

But with the emergence of AI agents, this approach will probably need to be deepened.

A classic chatbot mainly produces text.
An AI agent can call tools, modify data, interact with other systems, trigger workflows or produce real effects within an organisation.

The question therefore becomes:

Does the risk come only from the model, or from the power of action given to the system?

Let us take a simple example.

The same AI model can be used in three different environments:

Case 1 — Writing assistant
The model summarises internal notes.
Main risk: error, confusion, loss of time.

Case 2 — HR assistant
The model helps filter or rank job applications.
Main risk: discrimination, exclusion, harm to fundamental rights.

Case 3 — Agent connected to an ERP system
The model proposes or triggers financial actions.
Main risk: real economic effect, legal liability, financial loss or unauthorised data modification.

In all three cases, the model may be the same.
But the level of risk is not the same, because the execution environment is not the same.

This is why risk classification should not only ask:

Which model is being used?

It should also examine:

In what context does it act?
What data does it use?
What tools can it call?
What decisions does it influence?
What actions can it trigger?
What is its mandate?
Who validates?
What evidence is retained?
Who is responsible?

This distinction is essential for agentic systems.

With AI agents, risk becomes a property of execution. A system becomes more sensitive when it combines:

AI model

  • sensitive data
  • connected tools
  • autonomy
  • impact on people
  • insufficient human oversight
  • absence of usable evidence

This does not mean that the model has no risk of its own. General-purpose AI models may themselves be subject to specific obligations, especially when they present systemic risks. The European Commission explains that such risks may be linked to the most advanced models, but also to their reach, scalability or access to tools. European rules on general-purpose AI models also provide for transparency obligations and, for models presenting systemic risks, obligations to assess and mitigate those risks.

But for organisations deploying AI agents, the concrete risk often appears at the level of integration: accessible data, connected tools, autonomy granted, human supervision and the ability to demonstrate what was actually executed.

A future governance approach could therefore distinguish several layers of risk:

Model Risk
Capabilities, limitations and behaviours of the model.

Data Risk
Quality, sensitivity, provenance and freshness of the data.

Context Risk
Domain of use, affected population, business criticality.

Mandate Risk
What the agent is authorised to do, what is prohibited, duration and limits of the mandate.

Action Risk
Connected tools, triggerable actions, reversible or irreversible effects.

Human Oversight Risk
Real human supervision, competence of the supervisor, ability to intervene.

Evidence Risk
Ability to prove the intention, authorisation, action, result and responsibility.

This framework helps avoid a common mistake: believing that a more capable model is automatically safer.

A very advanced model, connected to critical tools without a clear mandate or execution proof, may be more risky than a more modest model used in a limited, supervised and auditable environment.

The central question therefore becomes:

Is an AI system risky because it is intelligent, or because it can act without sufficient governance?

This approach also aligns with risk management frameworks that emphasise the need to manage AI risks for individuals, organisations and society, while integrating trust into the design, development, use and evaluation of systems. The NIST AI Risk Management Framework, for example, is designed precisely to improve organisations’ ability to manage these risks in a structured way.

Over the coming years, Europe could play an important role in developing more operational classification methods, focused not only on the AI system itself, but also on its real use, context, autonomy, authority, power of action and execution evidence.

This could lead to an evolution in the debate:

AI Compliance
Compliance of the AI system.

Execution Governance
Governance of what the system actually does.

Evidence-Based AI Oversight
Ability to demonstrate intention, authorisation, action, result and responsibility.

In a world of autonomous agents, the question will no longer be only:

“Which model are you using?”

but:

“What was it authorised to do, in what context, under what supervision, with what impact, and with what evidence?”

This may be where the next stage of responsible AI lies: moving from regulation centred on systems to governance centred on real execution.

Etichete
AI Act AI Governance

Comentariis

Profile picture for user Abou Faoor Yasser
Trimis de Yasser Abou Faoor la Mie, 08/07/2026 - 18:12

I think the distinction between AI compliance and execution governance is becoming increasingly important.

One challenge we’ve been exploring is that “evidence” shouldn’t simply mean logs retained by the system itself. For agentic systems, execution evidence should be independently verifiable, tamper-evident, and preserve the complete decision context — not only the outcome, but also the authorization, approvals, policy evaluation, execution metadata, and responsibility associated with that action.

This shifts evidence from an operational record into something that can support audits, regulatory reviews, and cross-organizational trust without depending on the originating platform.


 


 

Profile picture for user Yanbolu Burhan
Trimis de Burhan Yanbolu la Vin, 10/07/2026 - 03:05

Jeremy, this is an excellent framing — and the layered risk model you propose (Model → Data → Context → Mandate → Action → Human Oversight → Evidence) maps directly to what we're building in production.

The last layer — Evidence Risk — is the one most frameworks quietly skip. Everyone agrees you need "ability to prove the intention, authorisation, action, result and responsibility." Almost no one ships infrastructure that actually produces that proof in a way a third party can verify independently.

That's what TBN Protocol does: every AI agent decision gets a cryptographically signed, timestamped receipt — verifiable by anyone without trusting the system that produced it, the vendor, or any intermediary. Not a log. Not a dashboard. A signed, independently checkable record.

Your question — "Is an AI system risky because it is intelligent, or because it can act without sufficient governance?" — has a practical corollary: governance without evidence is just a claim. The receipt is what turns governance from a description into a demonstrable fact.

We're registered under the EU AI Act (Article 50) and operational today with 1,000+ governed AI decisions publicly verifiable. Happy to discuss further — this is exactly the space we work in.

Profile picture for user Olujide Sheriff
Trimis de Sheriff Olujide la Mar, 14/07/2026 - 12:48

Excellent framing, Jeremy. I've seen first-hand how organisations can underestimate AI risk by focusing too heavily on the model itself, without fully considering the context of use and the potential consequences for people.

I completely agree that AI risk needs to be assessed across multiple layers, not just the model. Your proposed framework reflects the direction governance needs to take.

To answer questions such as "What was it authorised to do, in what context, under what supervision, with what impact, and with what evidence?", the recently proposed SAFR (Safeguards for Agentic Finance at Runtime) framework provides a practical example of how execution governance can be operationalised. It introduces a governance checkpoint between every agent decision and execution through four runtime components: Agent Identity, Controls Repository, Disposition Engine, and a tamper-evident Audit Log. Together, they ensure that no agentic action reaches execution without first being declared, authorised, assessed, and recorded.

I think this is where AI governance is heading beyond governing models to governing runtime execution, where authority, mandates, human oversight, evidence and accountability become first-class governance concerns.