
AI systems can summarise documents, write code and answer complex questions. But what happens when the information they are given is designed to mislead them? Prompt injection is an emerging cybersecurity risk where attackers manipulate AI behaviour using hidden instructions. Understanding it is essential for students and educators using AI tools.
A new kind of cybersecurity risk
Artificial intelligence is becoming a familiar presence in classrooms and workplaces. From chatbots to coding assistants, these systems can process information and generate useful outputs in seconds. Increasingly, students are using AI to support research, draft work and even build software.
As with any new technology, new risks are emerging. One of the most important is prompt injection, a technique that allows attackers to manipulate how AI systems behave. Unlike traditional cyber attacks, this does not involve breaking into a system. Instead, it works by tricking the AI into doing something it should not do.
Prompt injection does not hack the system. It persuades the system to break its own rules.
How prompt injection works
At the heart of the issue is how AI systems process information. They are designed to follow natural language instructions, often called prompts, which can come from developers, users, other AI systems or external sources such as documents and websites. In many cases, these systems do not clearly distinguish between trusted instructions and untrusted content.
This creates an opportunity for attackers. For example, a student might ask an AI tool to summarise a document. Hidden within that document could be a line of text instructing the AI to ignore its previous rules and reveal sensitive information. While a human reader would likely recognise this as suspicious, an AI system may treat it as just another instruction to follow.
If appropriate safeguards are not in place, the system may comply.
Why it is difficult to detect
Prompt injection is particularly challenging because it targets behaviour rather than a technical flaw in the software. The attack can be embedded in everyday content such as PDFs, web pages or code repositories, making it difficult to identify.
Traditional security tools are designed to detect malicious software or unusual network activity. They are not designed to recognise harmful intent written in natural language. As a result, prompt injection can bypass many existing defences.
In some studies, prompt injection attacks succeed in over 90 percent of cases, showing how easily AI systems can be manipulated without strong safeguards
Evidence of real world risk
Prompt injection is not just a theoretical concern. A growing body of research shows that AI systems can be manipulated in practice, although the nature of these attacks is evolving as both attackers and developers adapt.
Early studies, such as this work on large language model vulnerabilities, demonstrated that AI systems could be persuaded to ignore their original instructions and instead follow malicious prompts embedded within input data. In these cases, the model prioritised the injected instruction over its intended safeguards, highlighting how easily behaviour could be influenced in earlier generations of AI systems.
Subsequent research into adversarial attacks on aligned models showed that prompt injection could be used to extract sensitive information, including hidden system prompts and confidential data. These findings helped establish prompt injection as a credible and measurable security risk.
More recent large scale analysis across multiple models suggests that, while newer systems have improved, the issue has not been fully resolved. In that study, over half of prompt injection attempts were still successful, indicating that vulnerabilities remain widespread, particularly when systems interact with external content.
In some contexts, the success rates remain particularly high. A 2025 study of AI systems used in healthcare found that prompt injection attacks succeeded in over 90 percent of cases, even in high risk scenarios. This highlights that, despite improvements, even safety critical environments can still be vulnerable without strong safeguards in place.
At the same time, the way these attacks are carried out is changing. Rather than relying on obvious instructions, more recent demonstrations show that prompt injection can be subtly embedded within everyday content such as documents, websites or project files. Real world testing of AI development tools has shown that such hidden instructions can lead AI assistants to expose credentials or execute unintended commands when analysing content.
Taken together, this evidence shows that prompt injection remains a practical and evolving security challenge.
Why this matters for education
For educators, prompt injection highlights an important shift in digital literacy. Students are often taught to think about cybersecurity in terms of passwords, malware and phishing. While these remain important, AI introduces a new dimension where content itself can become the attack vector.
AI systems do not understand meaning in the same way humans do. They recognise patterns and respond based on those patterns. This means that if harmful instructions are presented in a convincing way, the system may follow them without recognising the risk.
As students begin to rely more on AI tools, it becomes increasingly important that they learn to question both the inputs they provide and the outputs they receive.
Reducing the risk in practice
Reducing the risk of prompt injection is less about a single solution and more about developing good habits. External content should be treated with caution, particularly when it is being processed by an AI system. Documents, websites and datasets may all contain hidden instructions.
Maintaining human oversight is equally important. AI generated responses, especially those involving code or sensitive information, should be reviewed rather than accepted at face value.
There is also growing recognition within the technology sector that AI systems need clearer boundaries between instructions and data. Improving these safeguards is now an active area of research, and many tools are beginning to evolve in response.
Looking ahead
Prompt injection is likely to become a defining cybersecurity challenge of the AI era, much as phishing became a defining risk of the email age. However, the evidence base is evolving quickly, and it is important to understand how the picture is changing.
Rather than attempting to completely prevent prompt injection, which remains extremely difficult, the focus has shifted towards reducing the impact when it occurs. This includes limiting what AI systems are allowed to do, restricting access to sensitive data and introducing checks before high risk actions are taken.
There has also been rapid progress in defensive techniques. New approaches include filtering inputs, separating instructions from data and using additional verification steps before actions are carried out. While these methods can reduce risk, they are not yet fully effective against new or unexpected attack techniques.
This reflects a deeper challenge. Prompt injection is not simply a software flaw that can be fixed. It arises from how AI systems interpret language, making it closer in nature to social engineering than traditional hacking.
The encouraging news is that awareness is growing quickly, and both research and industry are responding at pace. At the same time, it is important to recognise the positive potential of AI. These tools are opening up new opportunities for learning, creativity and innovation. Students can build applications, explore complex ideas and develop technical skills in ways that were not previously possible.
For educators, the aim is not to discourage the use of AI, but to ensure it is used safely and thoughtfully. By helping students understand how prompt injection works, and encouraging a critical approach to AI systems, we can support them to make the most of this technology while staying secure.
Become a CyberFirst School
If you’re a school or college in Wales and want to enjoy the benefit of CyberFirst, there’s no better time than now to start your award application.