What Is Shadow AI, and How Do You Get It Under Control?

Photo: Juan Pablo Serrano / Pexels
Key takeaways
- Shadow AI is any AI tool used at work that hasn't been approved or reviewed - often on a personal, free account.
- It's not a discipline problem: people adopt it because it's faster than waiting for approval that may never come.
- The real risk is confidential data entering a consumer tool's training pipeline, not the existence of the tool itself.
- Blanket bans tend to push usage further out of sight rather than stopping it.
- A fast approval path plus a living tool register shrinks shadow AI faster than any policy alone.
Shadow AI is the use of AI tools that your organisation hasn't approved, or in many cases doesn't even know exist: marketing running ChatGPT on a personal account, support installing a free AI browser extension, an engineer pasting a chunk of code into an assistant nobody vetted. It's the AI-era version of shadow IT, but it spreads faster, because there's often no purchase order, no install beyond a browser tab, and no IT ticket to catch it. If you've never audited what your team actually uses day to day, the honest assumption is that some shadow AI is already there.
Shadow AI vs shadow IT: what's actually different
Shadow IT usually means an unsanctioned app or spreadsheet tool: annoying, but rarely a data-training risk in itself. Shadow AI is riskier for a specific reason: many consumer AI tools don't just store what you give them, they can use it to train future versions of the model. That means data typed into a shadow AI tool doesn't just sit in one account waiting to be found in an audit. It can, depending on the tool and tier, become part of the product itself, in a way a rogue spreadsheet never could, and there's no easy way to claw it back once it has.
Why it happens
Because AI tools are useful, free, and a browser tab away, while formal approval is often slow or simply doesn't exist yet. People aren't being reckless. They're solving a real problem - a deadline, a repetitive task, a blank page - with whatever's available, and a proper procurement process is rarely available at 4pm on a Thursday. Banning AI outright doesn't remove that underlying pressure; it just moves the workaround further out of sight, onto personal devices where there's even less visibility than before. A fair bit of shadow AI, in our experience, comes from perfectly capable, well-meaning staff who simply had no approved option in front of them when they needed one.
Where the risk actually sits
The risk isn't the tool's existence. It's what goes through it once real data enters the picture. Consumer or free tiers of many AI tools train on user inputs by default unless someone actively opts out, and they may retain conversations for longer than most people assume. Confidential plans, customer records or proprietary source code typed into a free-tier tool can end up shaping a model that other users interact with later, or sitting exposed in a breach, with no record that any of it ever happened. Worse, the data doesn't need to be dramatic to matter: a draft contract, an unreleased pricing sheet or a client's home address in a support ticket is enough to turn a convenience into an incident.
Does an approved tool on the wrong tier still count as shadow AI?
Often, yes, and it's the subtlest version of the problem. A tool can be formally 'approved' at the brand level, say ChatGPT, while individual staff sign up on the free personal tier rather than the business tier that procurement actually vetted. The product name is identical; the data-handling terms underneath it are not. OpenAI's business tiers don't train on your data by default, but the free consumer tier can, unless data controls are switched off manually, and OpenAI's enterprise privacy documentation spells out exactly where that line sits. That gap between 'the tool is approved' and 'this specific account is safe' is where a lot of real exposure hides, and it's worth checking directly rather than assuming.
Retention windows are shrinking, but defaults still vary
The data-handling landscape is moving in the right direction industry-wide, but defaults still differ sharply between vendors, and even between tiers of the same vendor. Anthropic, for one, cut its API log retention from thirty days down to seven as of 14 September 2025, a genuine improvement, though that specific change applies to the API rather than to the free consumer Claude.ai product where most shadow-AI use actually happens. Anthropic's privacy centre sets out exactly which surfaces train on data and which don't. The honest takeaway is that these terms shift often enough that a one-off check isn't enough; it needs to sit inside a register someone actually revisits.
What doesn't work
Blanket bans read as decisive but rarely work. They don't remove the pressure that drove adoption in the first place, so usage doesn't stop, it just goes quiet: personal devices, personal accounts, no one mentioning it in a team meeting. A policy nobody can realistically follow produces silent non-compliance, not compliance, and it leaves you with less visibility than you had before you wrote the rule. Worth a gut-check here: if your current AI rule is simply 'don't use it,' assume it's already being ignored somewhere in the business.
How to control it without a ban
Give people an approved path that's genuinely as easy as the shadow one. Publish a short AI usage policy that names which tools and tiers are safe. Keep an approved-tools register that people can actually check before reaching for something new, and make it fast to request an addition - days, not months. When the sanctioned route is no slower than the shadow one, shadow AI shrinks on its own, because most people would rather not be the person who has to explain later why they went around the rules.
A quick example
A customer support lead at a thirty-person SaaS startup started using a free AI transcription tool to summarise calls, purely to save time between shifts. Nobody told her not to; nobody had told her anything, because no policy existed yet. It only came up when a customer asked, during a support call, whether their conversation was being processed by a third-party AI tool. It was, on the free tier, with unclear retention terms. The company's response wasn't to ban transcription tools; it was to approve one specific paid tier with a written data-processing agreement, and publish that choice so the next person didn't have to guess. The whole review, from the customer's question to a written answer everyone could point to, took under a week, mostly because nobody had to start from scratch.
Know your tools before you approve them
Start by checking how your most-used AI tools actually handle data: whether they train on it by default, how long they retain it, and whether a business tier exists that turns both off. Our AI Tool Risk Directory rates the popular ones - including ChatGPT, Claude and Microsoft 365 Copilot - from their own published policies, so you're not starting the review from a blank page. That single lookup, before a tool gets approved rather than after it's already in daily use, is usually the difference between shadow AI you catch early and shadow AI you only find out about from a customer.
| Signal | Where to look | What it usually means |
|---|---|---|
| Personal email domains signing up to AI tools | Expense claims, browser extension inventories | Likely a free or consumer tier, with no admin controls or DPA |
| A spike in traffic to AI-tool domains | SSO/OAuth app logs, firewall or DNS logs | An unapproved tool already in regular, unreviewed use |
| Customer or case data pasted into a chat window | DLP alerts, screen-share observations | Possible confidential data entering a training pipeline |
| 'Oh, we already use that' in a tool-approval discussion | Team surveys, onboarding conversations | Informal adoption that never reached a register or review |
“Shadow AI isn't a discipline problem, it's a service-gap problem: give people an approved tool that's as easy as the one they already found, and the shadow list gets shorter on its own.”