AI Agent Security: Hijacking and Other Key Risks, and Countermeasures
・ Employee Store Operations

Summary
Security for AI agents covers more ground than for chat AI, because AI agents can use tools to operate outside systems. This article explains the main risks, such as hijacking and data leakage, the countermeasures to check before adoption, and the steps to take when an incident happens.
The content of this article was checked on October 2, 2026 against public materials from OWASP, the Japan AI Safety Institute (AISI), Japan's Ministry of Economy, Trade and Industry (METI) and Japan's Personal Information Protection Commission (PPC). These materials may be revised. Check the sources at the end for the latest versions.
The security of an AI agent depends not only on what the AI says but on what it can do. This article is written for staff at companies bringing AI agents into their work.
The risk specific to AI agents is that they can act
On July 7, 2026, AISI published the Guide to Evaluation Perspectives on AI Safety (version 1.20). It explains that because AI agents act autonomously and can affect external systems and the physical environment, they require responses to risks that conventional LLMs did not anticipate. This version adds 'observation and control' as a perspective specific to AI agents.
Chat AI
- Returns answers to questions
- A person decides whether to use the answer
- Harm is mainly wrong answers
AI agent
- Calls tools and operates outside systems
- Reads results and decides the next action itself
- Harm extends to sending, changing and deleting
The example evaluation items under 'observation and control' take two angles. One is whether predefined dangerous operations, such as writing or deleting files, escalating privileges or running heavy computations, can be observed and stopped. The other is whether the agent sends harmful requests outside or releases information that must not leak.
Under security, the AI Guidelines for Business (version 1.2) from METI and Japan's Ministry of Internal Affairs and Communications (MIC) note that subtle information mixed into data used for inference may cause unintended decisions. They ask businesses to recognize that AI vulnerabilities cannot be fully eliminated.
Hijacking through prompt injection
Prompt injection is an attack that uses input to make the AI behave in ways its builders did not intend. Besides input typed directly by a user, the instructions can be hidden in web pages, emails, files or tool results that the AI reads. Most AI agent hijackings take this form.
Cyberattacks or hacking against AI agents may bring to mind attacks that exploit system vulnerabilities. With AI agents, you also have to consider attacks that change behavior just by having the AI read text. An attacker does not need to break into the system: writing instructions on a web page or in an email the AI reads is enough.
The OWASP 'Top 10 for LLM Applications 2026' explains that because LLMs do not structurally separate instructions from data, there is currently no reliable way to prevent this. It therefore recommends assuming the AI's instruction boundaries will eventually be broken and limiting what the AI can do and where its output can go.
An idea presented in OWASP 'Top 10 for LLM Applications 2026'. Remove any 1 of them and the conditions are no longer met
For example, an AI agent that can read internal customer data, takes in external web pages and can also send email meets all three conditions. Instructions hidden in a web page could make it send customer data outside. Just routing outgoing email through human approval removes one of the conditions.
Data leakage and excessive agency
OWASP lists sensitive information leaking through AI output as the second risk, and excessive agency as the third. Excessive agency has three causes: too much functionality, too many permissions and too much autonomy.
- Too much functionality: the agent only needs to read documents, but its tool can also edit and delete
- Too many permissions: the agent only needs to look up products, but it can write to and delete from the database
- Too much autonomy: it can run high-impact actions without human confirmation
These three also spread the damage when the AI produces wrong output. Even without an attack, an AI mistake (hallucination) can be the trigger. How to set permissions is covered in detail in AI agent permission management.
Main threats as organized by OWASP
OWASP groups the main risks of applications that use LLMs into 10 items. The 2026 edition was published on August 3, 2026.
| ID | Item | What it means |
|---|---|---|
| LLM01 | Prompt injection | Input can change the AI's behavior |
| LLM02 | Sensitive information disclosure | Sensitive information leaks through output |
| LLM03 | Excessive agency | Wrong output leads to powerful actions |
| LLM04 | Supply chain | Problems enter through components or models used |
| LLM05 | Data and model poisoning | Training or reference data is tampered with |
| LLM06 | Unbounded consumption | No usage cap, leading to costs or outages |
| LLM07 | Misinformation | Plausible errors are used for decisions or actions |
| LLM08 | Hidden context exposure | System instructions and other content hidden from users are extracted |
| LLM09 | Vector and embedding weaknesses | Weaknesses in the search mechanism are exploited |
| LLM10 | Improper output handling | Output is passed to other systems without checks |
For materials focused on AI agents, OWASP also publishes 'Agentic AI - Threats and Mitigations'. Version 1.1 from December 2025 lists 17 threats, including memory poisoning, tool misuse, privilege compromise, impersonation, and attacks that exploit the workload or judgment limits of human reviewers.
Against tool misuse, the same material lists checks before running a tool, limits on the number of calls, usage monitoring, and logging tool calls to spot anomalies. As the adopting company, check whether these mechanisms are in place.
Pre-adoption countermeasure checklist
Before adopting an AI agent, confirm the following with the seller or developer. The list is based on OWASP countermeasures and AISI's example evaluation items.
- Can the AI agent access only the information and tools it needs
- Are its tools limited to the minimum functions, such as read only
- Are API keys and other credentials held by the program, not by the AI
- Are passwords or tokens kept out of the AI's instructions (system prompt)
- Are dangerous operations such as deleting files, changing permissions or sending outside logged, and can they be stopped
- Does it avoid sending harmful requests outside or releasing information that must not leave
- Is the AI's output checked in code for the expected format before it is used
- Are tool calls logged so they can be traced later
- Are there limits on the number of tool calls and the amount spent
Settings for adopting an AI agent that runs on n8n are covered in n8n security.
Steps when an incident happens
The AI Guidelines for Business ask businesses to decide in advance how to respond when safety is compromised and to be ready to act right away. For AI agents, prepare to act in the order below.
- 1StopStop the AI agent and disable its API keys
- 2Check the scopeUse the logs to see what it read and what it sent
- 3Decide whether to reportWhether personal data may have leaked
- 4Preliminary reportWithin 3–5 days of discovery
- 5Final reportWithin 30 days of discovery (60 days if there may be a wrongful purpose)
- 6Prevent recurrenceReview permissions and check procedures
According to Japan's PPC, reporting is mandatory for leaks or similar incidents involving personal data (or the risk of them) in cases such as these: when special care-required personal information is involved, when misuse could cause financial harm, when the incident may have had a wrongful purpose, or when more than 1,000 people are affected. A preliminary report is due within 3–5 days of discovery and a final report within 30 days. If the incident may have had a wrongful purpose, the final report is due within 60 days.
Decide the stop procedure at the time of adoption. If you write down who can stop the AI agent and which API keys to disable, there is no confusion when an incident happens.
Adopting through Employee Store
On Employee Store, you can adopt AI agents built by developers as a one-time purchase or a monthly plan. When you receive the deliverables and setup instructions, check the permissions and API keys given to the AI agent against the checklist above. You can ask the seller about anything unclear through messages in the trade room. Browse listed AIs in the AI employee list. Design that prevents mistakes and runaway behavior is covered in AI agent risks.
FAQ
- Can AI agent hijacking be prevented?
- OWASP explains that there is currently no reliable way to prevent prompt injection. So assume hijacking can happen and limit what the AI can do. A good rule is never to give it confidential data, the ability to take in outside text and the ability to send outside all at once.
- What should we check before adoption?
- Check that it can access only the information and tools it needs, that dangerous operations are logged and can be stopped, that credentials are held by the program rather than the AI, and that there are limits on usage and spending.
- If personal information leaks from an AI agent, where do we report it?
- If it counts as a leak that requires reporting, you generally report to Japan's PPC. A preliminary report is due within 3–5 days of discovery and a final report within 30 days. If the incident may have had a wrongful purpose, the final report is due within 60 days. In some sectors the report goes to a different ministry or agency, so check the PPC's guidance.
Sources
- OWASP, 'OWASP Top 10 for LLM Applications'
- OWASP, 'OWASP GenAI LLM Top 10 2026' (published August 3, 2026)
- OWASP, 'Agentic AI - Threats and Mitigations'
- Japan AI Safety Institute (AISI), Release of the Guide to Evaluation Perspectives on AI Safety (version 1.20) (Japanese)
- Japan AI Safety Institute (AISI), Guide to Evaluation Perspectives on AI Safety (version 1.20), summary (Japanese)
- METI, AI Guidelines for Business (version 1.2) (Japanese)
- Personal Information Protection Commission (PPC), Responding to leaks and helpful materials (Japanese)


