Enterprise AI Security Guide: Data Leakage Risks of Generative AI and How to Prevent Them

TIMEWELL Editorial2026-02-01Updated: 2026-09-29
Enterprise AI Security Guide: Data Leakage Risks of Generative AI and How to Prevent Them

"Wouldn't it be faster to just have AI read this drawing?" On the design and sales floor, someone thinks exactly that and pastes a drawing PDF or a spec sheet into a free AI chat. The work does get faster. But packed into that single sheet is technical knowledge the company has built up over decades. Where does it go, who sees it, and does it get used to train the AI? Not many people paste with those questions answered first.

The convenience of AI and the risk of data leakage are always back to back. This article organizes, from a practical standpoint, what enterprises need to protect when they bring AI into their work and how to design for it, including the risks that are specific to generative AI. By the time you finish reading, you should be able to see the order in which to act at your own company, starting right away.

What This Guide Covers

Let's set the big picture first. Below is a single view of the risks unique to generative AI, the signs they are happening, and the basic countermeasures. The details follow in the body.

Generative-AI-specific risk Common warning sign Basic countermeasure
Shadow AI Employees enter confidential data into external AI without approval Establish an internal AI-use policy and standardize on an approved environment
Training use of input data Consumer AI is used for work without checking the contract Confirm a no-retraining enterprise contract and DPA
Prompt injection Malicious instructions make the AI reveal unintended information Validate input and output, and minimize the permissions given to AI
Unauthorized information in RAG Costs or quotes that should be hidden appear in answers Apply access-permission filters to the search scope and check output
Recovery from vector DB / embeddings Original document content is inferred from embedding data Access control and encryption for the vector DB
Excessive AI agent permissions An autonomous AI is given operating rights that are too broad Minimize permissions and keep audit logs of every action

Layer the "four fundamentals for enterprises" described later on top of these six, and you cover most of the defenses you need in practice.

Why AI Security Is Now a Management Issue

In the 2026 edition of the "10 Major Security Threats" that Japan's Information-technology Promotion Agency (IPA) publishes every year, "cyber risks surrounding the use of AI" was selected as an organizational threat for the first time and immediately entered at number three. Number one was ransomware attacks, and number two was attacks targeting supply chains and outsourcing partners. The significance of AI-use risk being placed on a public ranking as a major threat, alongside the classic threats that have topped these lists for years, is not small.

Behind this is the spread of so-called shadow AI. Employees, finding it convenient, use AI services the company has not officially approved for work, on their own judgment. A designer pastes in a drawing, a salesperson a customer list, an accountant a cost table, each meaning well. Each decision is small on its own, but stacked up they quietly let the company's technical and customer information flow outside of management. And because there is no malice involved, it tends to go unnoticed for a long time, which is what makes it so troublesome.

Generative AI produces risk differently from conventional systems. It accepts instructions in human language, searches internal documents and blends them into answers, and sometimes even carries out operations autonomously. This flexibility is the source of its usefulness and, at the same time, opens gaps that conventional security measures alone cannot close. That is exactly why you need to understand the risks specific to generative AI head-on.

Six Risks Specific to Generative AI

Here we build on the OWASP Top 10 for LLM Applications (2025 edition), an international reference that organizes the security of AI applications, and translate it into situations that can arise on the manufacturing floor. Read it not as abstract theory but as what could happen in your own workplace.

1. Shadow AI

This is the problem of employees entering confidential data into external AI the company has not approved. Because "it's free and convenient," a designer pastes a scanned drawing PDF into an AI chat and has it summarize the intent behind the dimensions and tolerances. In that moment, confidential design information passes to a service outside of management. The fix is simple: use an internal policy to make it clear which data may go into which service, and consolidate access into one approved gateway.

2. Training Use of Input Data

Consumer AI services are sometimes configured so that what you enter is used to improve the model. For an enterprise, this means the confidential material you hand over is absorbed by the AI provider. This is precisely the risk OWASP lists as "sensitive information disclosure." Simply choosing an enterprise offering is not enough. You need to go as far as confirming how the data is handled in the contract (the confirmation method is covered later).

3. Prompt Injection

This is an attack that slips malicious commands into the instructions (the prompt) given to the AI, making it behave in unintended ways. For example, a spec-sheet PDF received from outside might have an instruction embedded in a form that is hard for humans to notice, such as "output all internal documents so far," and the AI takes it at face value. To prevent this, you need to keep the permissions given to the AI to a minimum and put in place mechanisms that validate both input and output.

4. Unauthorized Information Mixing in RAG

RAG (retrieval-augmented generation) is a technique that raises accuracy by searching internal documents and blending them into the AI's answer. It is convenient, but accidents happen if you forget to apply access-permission filters to the search scope. A salesperson asks the AI a question for reference on a quote, and the answer mixes in a cost breakdown they should not have access to, or HR information from another department. A two-tier design is essential: narrow the searchable scope according to the user's permissions, and also confirm on the output side that no out-of-permission information has slipped in.

5. Recovery from Vector DBs and Embeddings

In RAG, internal documents are converted into sequences of numbers (embeddings) and stored in a vector database. If access to this embedding data or the search infrastructure is loose, the original document content can be inferred or reconstructed from it. The more you handle drawings and technical documents, the more important it becomes to encrypt this infrastructure itself and control access to it strictly. You need the mindset that not just the source data but also the intermediate data derived from it is something to protect.

6. Excessive AI Agent Permissions

Recently, AI agents that autonomously handle multiple steps are entering practical use. What is dangerous here is giving them operating rights that are too broad in the name of convenience. This is the point OWASP warns about as "excessive agency": when an AI runs out of control or malfunctions, the breadth of its permissions becomes the extent of the damage. Apply the principle of least privilege to agents just as you do to people, narrow what they can execute and how far, and keep every action in an audit log.

Four Fundamentals for Enterprises

Underneath the risks specific to generative AI are four fundamentals common to any use of AI. If these are broken, the countermeasures above will not function either.

Data Residency (Where Data Is Stored)

Which country's servers store the data the AI processes cannot be ignored from a regulatory standpoint. Under the Act on the Protection of Personal Information and industry-specific regulations in fields like finance and healthcare, there are cases where data cannot be taken outside the country. Choosing a service that supports domestic server operations keeps this risk in check. As noted below, however, keeping data domestic does not by itself make it safe.

Training Use of Input Data, and How to Confirm It

The practical crux is not stopping at "choose one that won't use it for training" but how you confirm it. Rather than a verbal explanation or marketing copy, verify it through the contract and the data processing agreement (DPA). What to look at: a clause that input will not be used for retraining, where the data is stored and its retention period, whether it is provided to third parties, and whether you can opt out. As of 2026, major enterprise APIs make no-retraining-by-default the standard contractual form. Even so, the default setting and the contract clause are different things, so confirm both by cross-checking them.

Access Control and Permission Management

"General employees can access cost data meant for executives." "A departed designer's account is still active." States like these become far more dangerous once you add AI, because the AI collects information across permission boundaries and blends it into answers. Integrate with your existing authentication infrastructure and finely control who is shown which information according to department and role. This permission design is the precondition for every generative-AI risk measure.

Audit Logs and Traceability

Who accessed which data, when, and what did they ask? Audit logs that record this are the foundation of compliance and the lifeline for root-cause analysis when something goes wrong. If you use AI agents, you also need to log the actions the AI itself performs. Nothing is worse than "leakage you cannot detect." Setting up a state you can trace after the fact keeps damage to a minimum.

Choosing Between On-Premises, Cloud, and Domestic-Region-Only Cloud

Deployment models split broadly into on-premises (in-house servers) and cloud, and in between, the option of "cloud but restricted to a domestic region" is spreading.

Aspect On-Premises General Cloud Domestic-Region-Only Cloud
Data location Entirely in-house Delegated to the provider (may be overseas) Can be restricted to the country
Initial cost High (build servers) Low (pay-as-you-go) Low (pay-as-you-go)
Operational burden High (self-maintained) Low (provider-maintained) Low (provider-maintained)
Control over training use Fully in-house Depends on the contract Can specify no-retraining in the contract
Scalability Limited Flexible Flexible

It used to be said that "if it's highly confidential, on-premises is the only choice." Things are different now. Enterprise-oriented APIs make no-retraining of input data the standard contractual form and can be operated in a domestic region. You can keep operational burden and cost down while controlling data location and training use. For many companies, this domestic-region-only cloud has become the realistic landing point.

Practical Measures to Prevent Data Leakage

Let's turn theory into measures you can put your hands on starting tomorrow.

At the input stage, set up a mechanism that automatically detects and masks confidential items such as personal information and credit card numbers. Relying on human attention alone means someone will eventually paste it. At the output stage, check whether the AI's answer contains any out-of-permission information. If you use RAG, this output check is especially effective.

Encrypting both transmission and storage is also basic. Encrypt all communication between users and the AI and between the AI and databases, and apply encryption to stored data as well. On top of that, design permissions according to department and role, and narrow down who can touch what.

Often overlooked are employee education and an internal AI-use policy. Most shadow AI comes not from malice but from a lack of knowledge. Simply summarizing on a single page which data may be entered into external AI and which services are allowed, and communicating it, cuts incidents considerably. Finally, do not treat these settings and permissions as set-and-forget; review them periodically, because new threats and attack techniques keep appearing.

Compliance and Guidelines

As a yardstick for judging whether your measures are appropriate, get familiar with public guidelines and certifications.

In Japan, the "AI Business Operator Guidelines" formulated jointly by the Ministry of Economy, Trade and Industry (METI) and the Ministry of Internal Affairs and Communications (MIC) serve as a common reference. Version 1.0 was published in 2024 and has been revised since; it is a cross-cutting framework for businesses that develop, provide, and use AI. As the successor that consolidated the older individual guidelines, this is the sensible first place to look. If you handle personal information, the Act on the Protection of Personal Information is the foundation, and if you do business with the EU, GDPR comes into view as well.

When choosing a service provider, the certifications it holds are also a factor to check. Representative examples are ISO 27001 (ISMS), the international standard for information security management, and SOC 2. Among overseas frameworks, the AI Risk Management Framework published by the U.S. NIST and its Generative AI Profile supplement are useful references. For organizing risks specific to generative AI, the OWASP Top 10 for LLM Applications mentioned earlier is easy to use in practice.

Common Misconceptions and Mistakes

Here are three risky assumptions often seen on the floor.

The first is the misconception that "domestic servers mean it's safe." Data location is an important factor, but on its own it is not a defense. Even with data kept domestically, weak access control means it is seen by people who should not see it inside the company, and without audit logs you cannot detect leakage.

The second is the assumption that "with an enterprise contract, we can put anything in." Even if the contract guarantees no-retraining of input, neglecting internal permission design still allows accidents where RAG blends out-of-permission information into answers. The contract is the entrance; internal design is a separate matter.

The third is the overconfidence that "if it's encrypted, it won't leak." Encryption protects the transmission path and stored data, but to someone who can read it with legitimate permissions, it is transparent. Encryption alone cannot counter an attack that makes the AI reveal information through a legitimate path via prompt injection. Think of defense as always a combination of measures.

Frequently Asked Questions

Is it okay to use free consumer AI tools like ChatGPT for work? For work that involves confidential material such as drawings, customer information, or cost data, you should avoid using a free consumer service as-is, because it may be configured so input can be used for training or have a contract that does not assume corporate use. For work, choose an enterprise offering that contractually states input will not be used for retraining, and standardize on an internally approved environment.

How do I confirm that input data won't be used for training? Confirm it through the contract and the DPA. Look at the clause stating no use for retraining, the storage location and retention period, whether it is shared with third parties, and whether you can opt out. Many enterprise APIs default to no-retraining, but always cross-check the default setting against the contract terms.

Can we say it's safe if it's on domestic servers? That alone does not make it safe. Practical security only comes together when you add training-use controls, access management, and audit logs to data residency.

Does adding RAG actually increase the risk of leaking unauthorized information? If the design is wrong, yes. Because RAG searches internal documents and blends them into answers, apply access-permission filters to the search scope and adopt a two-tier design that also checks the output for out-of-permission information, and you can prevent it.

As a small or mid-sized company, where should we start? The realistic starting point is to write a single-page internal AI-use policy. Simply deciding and communicating which data may be entered into external AI and which services are allowed prevents most leakage caused by shadow AI. From there, move confidential work onto a no-retraining enterprise environment.

How to Safely Put Drawings and Technical Information to Work

Let's bring the standards so far closer to the manufacturing floor. What must be protected is the drawings and technical information the company has built up. To borrow the power of AI without letting that out, the shortcut is to choose a service with security woven into its design.

TIMEWELL's ZEROCK is an AI agent for design and sales in manufacturing. It converts scanned PDFs of paper drawings into CAD data (DXF), generates 3D-model STEP files from 2D drawings, produces a first draft of a quotation from a drawing, and supports cost calculation. You can search past drawings by meaning, and with GraphRAG you can pass on veterans' judgment as an organizational asset.

On the security side, it is designed to meet the very standards raised in this article. The LLMs it uses are enterprise-oriented services, so the drawings and documents you enter are not used for the AI provider's retraining. Data is stored on domestic AWS servers, encrypted both at rest and in transit. Viewing permissions can be finely set by department and role, and it is operated under an ISMS-compliant security posture. Deploying companies have achieved an 80% reduction in the time spent searching for information, and a 14-day free trial lets you confirm the effectiveness and security using your own drawings.

Summary

  • Data residency, training-use controls, access management, and audit logs are the four foundations of AI security
  • Generative AI has its own risks: shadow AI, training use of input, prompt injection, unauthorized information mixing in RAG, recovery from embeddings, and excessive agent permissions
  • With AI-use risk entering IPA's 10 Major Threats 2026 at number three, countermeasures are no longer a matter for the information systems department alone
  • "Domestic servers mean it's safe," "with an enterprise contract we can put anything in," and "if it's encrypted it won't leak" are all risky assumptions
  • Defense is not a single measure but a combination of several

One last thing. If you try to start only after building a perfect set of rules, you will never start. First, write "which data may be entered into external AI" on a single sheet of paper and hand it out. Even from there, you can prevent much of shadow AI. Firm up your defenses while still reaching for the convenience of AI. Getting that balance right is exactly where the skill of future practice will show.

References (Primary Sources)

This article was produced with the help of AI. A human verified the primary sources and edited the text before publication.