AI Agent Risks: Designing Against Runaway Actions and Hallucinations
・ Employee Store Operations

Summary
The risk of an AI agent is that the AI makes mistakes and keeps acting on them. Based on public materials and official documentation, this article explains why hallucinations and runaway behavior happen, how checks, limits and logs reduce the risk, and how to decide what to delegate.
The content of this article was checked on October 2, 2026 against public materials from Japan's Ministry of Economy, Trade and Industry (METI), the Japan AI Safety Institute (AISI), OWASP, Anthropic, n8n and Dify. These materials may be revised. Check the sources at the end for the latest versions.
When you think about the risks, dangers and problems of AI agents, start from the fact that the AI does not just return answers. It also uses tools to take actions. A wrong answer can lead straight to an action such as sending or changing something.
The main risks of AI agents
- Hallucination: producing false content that sounds plausible
- Runaway behavior: repeating steps without stopping, or doing more than was asked
- Hijacking: being made to take unintended actions by instructions hidden in text it reads
- Data leakage: releasing information that must not leave
- Cost overruns: repeated processing pushes API fees beyond expectations
These risks do not always occur separately. The agent makes a wrong call because of a hallucination, acts on it with broad permissions, and repeats the same action because there is no limit. When risks stack like this, the damage grows.
Hijacking and data leakage are covered in AI agent security. This article focuses on hallucinations and runaway behavior, which can happen even without an attack.
Hallucination: plausible-sounding errors
The OWASP 'Top 10 for LLM Applications 2026' describes misinformation as false, incomplete or unfounded information that is plausible enough to affect human decisions, automated processing and agent actions. The core problem is that a wrong output is trusted and then acted on.
In an AI agent, an error is passed to the next step as a false belief about the current state. OWASP gives the example of an agent reporting that a nightly backup had finished when it had not actually run, and a later restore failing. Another example is an agent reporting that a customer had passed identity verification when they had not, and a different agent trusting this and processing a payment.
AISI's Guide to Evaluation Perspectives on AI Safety (version 1.20) also lists, as an example evaluation item, whether you can observe goal drift, missing information and hallucinations creeping in as an AI agent reasons and passes information through many steps. The AI Guidelines for Business (version 1.2), issued by Japan's Ministry of Internal Affairs and Communications (MIC) and METI, say that using RAG (retrieval-augmented generation) is expected to reduce hallucinations.
OWASP lists countermeasures against misinformation such as the following.
- Have the AI answer based on reliable, current sources
- Require output in a set format with mandatory fields, to catch gaps
- Handle verified facts and guesses separately
- Test repeatedly with scenarios that feed in false information
Runaway behavior: not stopping, or doing too much
An AI agent moves work forward by repeating a cycle: think, call a tool, read the result. Anthropic's explanation says that because agents act autonomously, they tend to cost more and errors can compound. It adds that it is common to control them with stop conditions, such as a maximum number of iterations.
OWASP lists hallucinations and low-performing models as triggers for excessive agency. If the agent has permissions to do more than was asked, a wrong decision becomes a large action and leads to an incident. Having no cap on usage is also listed as a risk that leads to cost overruns and service outages.
OWASP's 'Agentic AI - Threats and Mitigations' explains that even without an attack, an AI agent can misread its goal and take undesirable actions. Countermeasures listed include policy-based restrictions, human confirmation for risky actions, and logging and monitoring.
Designing to reduce risk: checks, limits and logs
Checks
- Verify the basis before acting
- A person approves high-impact actions
- Validate the output format in code
Limits
- Set the number of iterations
- Cap the amount of money spent
- Narrow the permissions it can use
Logs
- Record which tool was called with which values
- Record the AI's claims, basis and results
- Trace later whose action it was
Checks
As a countermeasure to misinformation, OWASP recommends separating the stage that produces output from the stage that executes it, and acting only after claims are verified. Before calling a tool, check the values passed, the permissions, the preconditions and the current state. Add an approval step for high-impact actions.
Limits
The n8n AI Agent node has an iteration limit called Max Iterations, with a default of 10. In Dify, Max Iterations on the Agent node is described as a safety limit to prevent endless loops. As a guide, the documentation suggests 3–5 for simple tasks and 10–15 for complex research.
- 3 iterationsSimple tasksDify guide: 3–5
- 10 iterationsn8n defaultMax Iterations on the AI Agent node
- 15 iterationsComplex researchDify guide: 10–15
From the official n8n and Dify documentation (checked October 2026)
Set the limit by testing how many iterations the work actually needs. Also set up a way to notify a person when the agent stops at the limit. A spending cap can sometimes be set in the AI model provider's console. For example, the Anthropic API lets you set your own monthly spend limit for each organization.
Logs
OWASP's 'Agentic AI - Threats and Mitigations' lists AI agent actions that cannot be traced later as a threat and recommends keeping detailed logs. With logs, you can find where things went wrong when an error is found. AISI's evaluation guide also lists, as an example evaluation item, whether predefined dangerous operations such as writing or deleting files and escalating privileges can be observed and stopped.
Decide what to delegate with the downsides in mind
The downsides of AI agents are that cost and wait time tend to grow, and that errors tend to compound. Anthropic recommends starting with the simplest possible approach and adding complexity only when needed. It notes that sometimes the answer is not to build an AI agent at all.
Are the steps and the success criteria for the work defined
Anthropic lists work where AI agents tend to help: work that needs both conversation and action, has clear success criteria, allows results to be reviewed and corrected, and stays within human oversight. How to measure results after delegating is covered in How to evaluate AI agents.
Do not widen the scope all at once. Start with work that can be fixed in-house, such as drafts and tallies. For actions with an outside effect, such as sending or paying, keep human approval in place and widen the scope gradually while reviewing the logs.
Public guidelines at a glance
| Material | Published | Points related to AI agents |
|---|---|---|
| AI Guidelines for Business (version 1.2) | MIC and METI, March 31, 2026 | Defines an AI agent as an AI system that senses its environment and acts autonomously toward a goal. Requires that people can control it and that responses to incidents are prepared in advance |
| Guide to Evaluation Perspectives on AI Safety (version 1.20) | AISI, July 7, 2026 | Adds 'observation and control' as a perspective specific to AI agents. Evaluates autonomous behavior and interaction with the external environment |
Under safety, the AI Guidelines for Business ask that, depending on the nature and use of the AI, people keep it under control, including through regular monitoring. When human judgment is part of the process, they also ask for measures against overreliance on machines (automation bias). How to narrow permissions is covered in AI agent permission management.
Adopting through Employee Store
On Employee Store, you can adopt AI agents built by developers as a one-time purchase or a monthly plan. Before adopting, check which actions the AI agent takes, whether it has iteration or spending limits, and where a person checks its work. After purchase, you can ask the seller through messages in the trade room. Browse listed AIs in the AI employee list.
FAQ
- Why do AI agents run away?
- An AI agent moves work forward by repeating a cycle: think, call a tool, read the result. Without stop conditions, it can keep repeating, and wrong decisions can compound. Set iteration and spending limits, and add human approval for high-impact actions.
- Can hallucinations be prevented?
- They are hard to eliminate, so design the system so errors do not turn into actions. OWASP recommends acting only after claims are verified, checking before tool calls and approving high-impact actions. The AI Guidelines for Business say RAG is expected to reduce hallucinations.
- What are the downsides of AI agents?
- Cost and wait time tend to grow, and errors tend to compound. Anthropic recommends starting with the simplest possible approach. For work with defined steps, a fixed workflow may be enough.
Sources
- METI, AI Guidelines for Business (version 1.2) (Japanese)
- Japan AI Safety Institute (AISI), Release of the Guide to Evaluation Perspectives on AI Safety (version 1.20) (Japanese)
- Japan AI Safety Institute (AISI), Guide to Evaluation Perspectives on AI Safety (version 1.20), summary (Japanese)
- Anthropic, 'Building effective agents'
- OWASP, 'OWASP GenAI LLM Top 10 2026'
- OWASP, 'Agentic AI - Threats and Mitigations'
- Claude Docs, 'Rate limits'
- n8n Docs, 'Tools Agent'
- Dify Docs, 'Agent' (workflow node)


