
AI systems are beginning to do more than answer questions or generate content. Increasingly, they can take action by sending emails, accessing files and interacting with software tools. These systems, known as AI agents, create new opportunities but also introduce new cybersecurity risks. Understanding the security risks of AI Agents is essential for teachers and students as AI becomes more integrated into everyday learning and work.
A shift from thinking to acting
Artificial intelligence is evolving quickly. Many people are already familiar with tools that can generate text, write code or summarise information. However, a new generation of systems is emerging that can go further.
These systems, often called AI agents, are designed not just to respond to prompts but to take action on behalf of a user. This might include searching the web, sending messages, updating documents or interacting with other software.
This shift is significant. When AI moves from generating ideas to carrying out tasks, the potential impact of mistakes or manipulation becomes much greater.
When AI moves from thinking to acting, the risks move from digital outputs to real world consequences.
What are AI agents?
AI agents combine language models with access to tools, data and systems. Instead of providing a single response, they can plan and carry out a sequence of actions to achieve a goal.
For example, an AI agent might begin by gathering information from a document, then use that information to draft a message, and finally send that message through an email system. In other cases, an agent may access online services, update files or coordinate tasks across multiple platforms.
Research into autonomous AI systems shows that these agents can break down complex goals into smaller steps and execute them with limited human input, as explored in recent work on tools using language models.
This ability to act independently is what makes AI agents powerful, but it is also what introduces new types of risk.
How Security risks of AI aGents CAN increase
The key difference with AI agents is that they interact directly with real systems. This creates a clear link between AI behaviour and real world outcomes.
When an AI system produces an incorrect answer, the impact may be limited. When an AI agent takes action based on incorrect or manipulated information, the consequences can be more serious. This might involve sending sensitive information to the wrong person, making unintended changes to documents or interacting with external systems in ways that were not expected. Further research on agent based AI systems has shown that when models are given access to tools, they can be influenced to perform actions that were not intended by the user. This highlights how the addition of real world capabilities changes the nature of the risk.
A growing concern is that AI systems are starting to reduce the barrier to creating real cyberattacks. When models become capable of identifying vulnerabilities and helping assemble working exploits, skills that once required specialist knowledge can become easier to access. This “democratisation” of exploit development changes the threat landscape: more people may be able to attempt sophisticated attacks, more quickly, and at larger scale. This is one reason why some advanced AI security capabilities are being tested with defenders first and released in a more controlled way to give organisations time to patch weaknesses before similar tools become widespread.
The link to prompt injection
AI agents are particularly vulnerable to the type of attacks discussed in the previous article. Prompt injection becomes more significant when an AI system is able to act.
When AI agents are connected to tools, a simple prompt injection attack can lead to real actions, not just incorrect answers.
If malicious instructions are hidden within content and an AI agent processes that content, the system may interpret those instructions as legitimate. It may then follow them, accessing data or performing actions without the user realising what has happened.
This research into indirect prompt injection has shown how hidden instructions in web pages or documents can influence AI systems in this way. When these systems are connected to tools, the effects can move beyond incorrect outputs and into real actions.
Trust when AI Systems collaborate on our behalf
This is a fundamental shift from understanding to experiencing trust.
Trust in technology is something most people rarely think about. It has been built gradually through everyday experiences.
For example, when someone opens an online shopping app, they trust that the products they see are accurate, that payments will be processed securely and that orders will arrive as expected. The system feels reliable because the process is familiar. Even if the technology behind it is complex, the steps are visible and the user remains in control.
This is an example of human to system trust. It is based on understanding, experience and the ability to intervene when something does not seem right.
AI agents represent a significant step beyond this.
Instead of simply helping users complete tasks, AI agents can now carry them out. More importantly, they do not operate alone. They are designed to work with other systems, and increasingly, with other AI agents.
For example, a single request such as organising a trip, managing a task or processing information may involve multiple systems working together, including the sharing of personal information or the transfer of funds. One AI agent may gather information, another may analyse it, and another may take action such as booking, updating or communicating. Each system relies on the output of the previous one.

From the user’s perspective, this appears as a single, seamless interaction. In reality, it is a chain of decisions and actions carried out across multiple systems, often without direct human involvement.
This is a major shift. Technology is no longer just something we use. It is something that acts and collaborates on our behalf.
The rise of transitive trust
The user is no longer just trusting one system. They are trusting a network of systems, including systems they may not know exist. Each AI agent in that chain must decide whether to trust the next system it interacts with.
This introduces what is known as transitive trust. If one AI agent trusts another, and that agent relies on a third, then trust is passed along the chain. The original user does not make these trust decisions directly, but they are still affected by them.
When AI agents work together on your behalf, trust is no longer a single choice. It becomes a chain reaction you do not see.
This has important implications for cybersecurity.
In traditional systems, trust was largely contained within a single interaction. In an environment where AI agents collaborate, a weakness in one system can influence many others. An error, a manipulation or a compromised service can be passed along the chain, affecting outcomes in ways that are difficult to trace.
The challenge is no longer just whether a system can be trusted, but whether the systems it trusts, and the systems they trust in turn, can also be relied upon.
Evidence from real world systems
Evidence from both research and industry shows that these risks are already emerging.
Studies of AI agents interacting with external tools have demonstrated that systems can be manipulated into performing unintended actions, including accessing or exposing sensitive information. These findings show that the issue is not limited to theory, but can occur in practical scenarios.
This makes the use of AI agents not just a technical issue, but a matter of digital literacy and safeguarding.
Industry reports also point to similar risks. In some cases, AI systems integrated into everyday tools have responded to hidden instructions in content such as calendar invites or documents, leading to unintended behaviour. These examples highlight how easily AI agents can be influenced when they operate across multiple systems.
Why this matters for education
For teachers and students, AI agents represent both an opportunity and a responsibility.
These systems have the potential to support learning by helping students organise information, manage tasks and explore ideas more efficiently. They can also reduce the time spent on routine activities, allowing more focus on understanding and creativity.
At the same time, they introduce new risks that need to be understood. Students need to recognise that AI systems can act on their behalf, and that those actions may involve real data and real systems. They also need to understand that not all information processed by AI can be trusted.
For teachers, there are additional considerations. As AI tools become more integrated into educational settings, it is important to think carefully about how they interact with student data, school systems and external platforms.
Reducing the Security risks of AI Agents in practice
Managing the risks of AI agents is not about avoiding the technology, but about using it carefully and thoughtfully.
One of the most effective approaches is to limit what AI systems are allowed to do. This includes restricting access to sensitive data and ensuring that systems only have the permissions they need to carry out specific tasks.
Human oversight also remains essential. Even when AI systems can act independently, their actions should be reviewed, particularly in educational contexts where accuracy and safety are important.
There is also growing emphasis on designing systems so that the impact of any single error is limited. This includes separating instructions from data and introducing checks before actions are carried out. These approaches do not remove the risk entirely, but they help to reduce its potential impact.
Looking Ahead
AI agents are likely to become a common part of both education and the workplace. As they become more capable, they will also become more embedded in everyday systems.
The evidence suggests that while these systems offer significant benefits, they also introduce new forms of cybersecurity risk. These risks are closely linked to earlier challenges such as prompt injection, but with more direct and immediate consequences.
At the same time, progress is being made. AI companies are developing new approaches to limit what agents can do and to reduce the impact of potential attacks. This includes better control over permissions, improved monitoring and safer system design.
However, the underlying challenge remains. AI agents operate by interpreting language and context, which makes them fundamentally different from traditional software systems.
For education, this presents an opportunity. By helping students understand how AI agents work, and by encouraging critical and responsible use, we can ensure that these tools are used safely and effectively.
Used well, AI agents have the potential to support learning, creativity and productivity. The key is to ensure that this potential is matched with awareness, understanding and good practice.
Become a CyberFirst School
If you’re a school or college in Wales and want to enjoy the benefit of CyberFirst, there’s no better time than now to start your award application.