Generative AI at Work: Precautions on Data Leaks, Copyright and Privacy
・ Employee Store Operations

Summary
When you use generative AI at work, the precautions depend on what you input and how you use what comes out. This article draws on materials from Japan's Personal Information Protection Commission, the Agency for Cultural Affairs, MIC and METI, plus each service's official pages, and ends with a short rule set for employees.
This article was checked against official pages from each agency and company on October 2, 2026. Service settings and terms can change. Before use, check the latest information through the sources at the end.
The main precautions for using generative AI at work fall into three areas: data leaks, personal information and copyright. The thinking is the same whether you use a chat-based generative AI or an AI agent that runs work automatically. Because an AI agent reads and writes data on its own, people tend to lose track of what is being input. Set the scope in your rules first.
Three common problems in business use
- Data leaks: confidential information you input is used by the service provider, for example for training
- Personal information: entering customer or employee personal data without checking consent or the scope of the purpose of use
- Copyright: the output resembles an existing work and is published outside the company as is
The AI Guidelines for Business from Japan's Ministry of Internal Affairs and Communications (MIC) and Ministry of Economy, Trade and Industry (METI) also ask businesses that use AI to take care not to input personal or confidential information inappropriately. General generative AI risks are covered in AI agent risks.
Beyond these three, watch for errors in the output. Generative AI responses can contain inaccurate content. The guidelines ask businesses that use AI to understand the accuracy of output and the level of risk before using it. Check numbers, dates and the names of people and companies against the original material before use.
Information you may and may not input
Whether to input something depends on two things: the type of information and the settings of the service you use. The same information may be fine to input into a business service set not to use data for training. On the other hand, it is safer not to put confidential information into a service whose settings you have not checked.
Do not input
- Confidential information into a service set to use data for training
- Personal information outside the purpose of use
- Passwords or API keys
Input after checking conditions
- Customer or employee personal data
- Internal-only documents
- Text or images created by others
OK to input
- Your company's already public information
- Data changed so individuals cannot be identified
- Drafts you wrote yourself
The appendix to the AI Guidelines for Business says that if confidential information entered into a generative AI service is going to be used by the provider as training data, you should take care not to input prompts containing it. For information in the conditional column, input it only after the checks in the following sections.
The same appendix also says not to give confidential information to AI carelessly, for example by becoming too emotionally attached to it. When you use AI in a conversational way, you can end up writing internal-only details while treating it like a consultation. Users are also expected to notify the service provider when they have security concerns.
The Personal Information Protection Commission's alert
On June 2, 2023, Japan's Personal Information Protection Commission (PPC) issued an alert on the use of generative AI services. For businesses that handle personal information, it raises two points.
- When entering prompts that contain personal information, confirm carefully that this is within the scope needed to achieve the specified purpose of use
- Entering prompts that contain personal data without the person's consent, where that data is then handled for purposes other than producing the response, may violate the Act on the Protection of Personal Information (Japan). If you enter such data, confirm carefully that the provider does not use it for machine learning, among other things
For general users, the alert notes that personal information you input may be used for machine learning and then output inaccurately. Responses may also contain inaccurate personal information. It asks users to decide after checking the provider's terms of use and privacy policy.
When you adopt an AI agent that handles customer names or inquiry details, these two points become your checklist: is it within the purpose of use, and does the provider refrain from using the data for training.
On the same day, the PPC also issued an alert to OpenAI, which develops and provides ChatGPT. It concerned the collection of sensitive personal information and notice of the purpose of use, among other things. The commission said it may issue further alerts. After you write your internal rules, check the commission's notices from time to time.
Copyright: the Agency for Cultural Affairs' view
On March 15, 2024, Japan's Agency for Cultural Affairs compiled its “General Understanding on AI and Copyright.” On July 31 of the same year, it published a checklist and guidance on AI and copyright that organizes measures by role. Companies that use AI at work should read the section for AI users (business users).
- 1Check the terms of useThey may prohibit inputting others' works
- 2Check the purpose of the inputInput meant to imitate an existing work may need permission
- 3Check for similaritySearch the internet, for example
- 4Keep a record of how it was generatedRecord the prompts used and similar details
- 5Use it outside the company
Based on the business user items in the Agency for Cultural Affairs' checklist and guidance on AI and copyright
The checklist explains that copyright infringement requires both similarity to an existing work and reliance on it. For that reason, it says you first need to check that the output does not resemble existing works before using it. As example methods, it lists internet text searches and image searches.
Note that generating something and using what you generated are judged differently. Generating output only for internal review may be lawful within the scope of copyright limitation provisions. Using it outside, such as distributing it online, is often said to fall outside that scope.
The guidance also explains that putting the title or character names of an existing work into a prompt can suggest that you knew of that work. Teaching employees the basics of copyright so they do not misunderstand is another measure the checklist lists.
The checklist also asks companies to take measures to prevent infringement, such as internal rules on how to use generative AI. It recommends deciding in advance how to respond if infringement does occur. The listed steps are sharing information, stopping and restoring AI use, dealing with the rights holder, finding the cause and preventing recurrence.
Check each service's training settings
Whether your input is used for training differs by service and contract type. Check the official explanations of the main services.
Which service and contract do you use?
OpenAI
OpenAI's page on how data is used explains that for consumer services, it may use your content to train models. If you turn this off in the privacy portal or data controls, new conversations are no longer used for training. Its enterprise privacy page explains that business data from ChatGPT Business, Enterprise, the API and similar offerings is not used for training by default.
Not being used for training is different from not being stored. According to the same page, in ChatGPT Business and Enterprise, workspace admins can set data retention periods. Deleted conversations are removed from OpenAI's systems within 30 days, unless retention is required by law or in similar cases. API inputs and outputs may be retained for up to 30 days to provide the service and detect abuse.
Handling confidential information in Dify
The privacy policy of LangGenius, which operates Dify, says it does not train AI models itself and does not use data from AI interactions for training. However, if you choose your own model provider and API key, data is sent directly to that provider. Whether it is used for training depends on the contract between you and that provider.
If you handle confidential information in Dify, check the terms of the model provider you connect as well as Dify's own settings. Even if you run Dify on your own server, data goes to the provider whenever you call a model through an external API.
A short rule set for employees
Here is an example that condenses the points above into something employees can remember.
- Use only services and accounts the company has approved
- Do not input passwords, API keys or internal-only documents. Consult the person in charge when needed
- Input customer or employee personal data only within the purpose of use, and only into services set not to use data for training
- Check the facts before using the output. Pay special attention to numbers and proper nouns
- Before publishing generated text or images outside the company, search to check they do not resemble existing works
- Keep a record of anything published outside the company, together with the prompts used
- If you notice odd behavior or a leak, stop using it at once and report to the person in charge
Company-wide governance is explained in AI agent governance, following Japan's national guidelines.
Extra checks when adopting an AI agent
An AI agent reads data from connected systems even without a person typing anything. In addition to the chat precautions, check the following before adoption.
- Which data in which systems it reads and writes
- Which AI model provider the data is sent to, and that provider's terms on training use
- Whether it runs in the provider's cloud or in your own environment
- Where input and output records are kept, and who can see them
On Employee Store, you can adopt AI employees with a one-time purchase or a monthly plan. Buyers and sellers can message each other in a deal room created for each transaction. Confirm the points above with the seller in the deal room before adoption. The basics of AI employees are explained in What is an AI employee.
FAQ
- Can employees use generative AI with personal accounts?
- We recommend that the company decide which services and accounts may be used. OpenAI explains that consumer services may use input for training, while business services do not use it for training by default. Handling changes with settings and contracts.
- Is it safe to put internal confidential information into Dify?
- Dify's operator says it does not use data from AI interactions for training. However, if you use a model with your own API key, data is sent to that model provider, and whether it is used for training depends on your contract with the provider. Check the terms of the provider you connect as well.
- Can images or text made with generative AI be published outside the company as is?
- The Agency for Cultural Affairs' checklist says you need to check that output does not resemble existing works before using it. Check with internet text and image searches before use.
Sources
- Personal Information Protection Commission, Alert on the use of generative AI services (checked October 2, 2026) (Japanese)
- Personal Information Protection Commission, Alert on the use of generative AI services (Attachment 1) (Japanese)
- Agency for Cultural Affairs, AI and Copyright (Japanese)
- Agency for Cultural Affairs, Checklist and Guidance on AI and Copyright (Japanese)
- MIC and METI, AI Guidelines for Business (version 1.2), appendix (Japanese)
- METI, AI Guidelines for Business (page for earlier versions) (Japanese)
- OpenAI, How your data is used to improve model performance (checked October 2, 2026)
- OpenAI, Enterprise privacy (checked October 2, 2026)
- LangGenius, Privacy Policy (Dify, checked October 2, 2026)


