Article
Your IP Is Leaking Into the Model
A critic and an insider are both telling you the same thing about where your proprietary research goes when someone pastes it into the wrong AI tool. Neither of them had to.
Compliance & Governance · Charles Samuel · 6 MIN READ · July 20, 2026
Contents
The warning a critic and an insider are both giving
Alex Karp and Satya Nadella agree on almost nothing. The former helms Palantir and has spent the past year publicly needling the rest of the AI industry for selling hype instead of results. Nadella helms Microsoft, which owns a meaningful stake in the company most responsible for said hype. One of them has every incentive to talk his rivals down. The other has every incentive to talk his own product up.
They landed on the same warning anyway.
Karp told CNBC that what enterprises actually want from an AI partnership is control over their own compute, models, data stack, and edge—not a vendor relationship where the vendor quietly learns everything you know. Nadella, on a separate podcast about how companies compound advantage, put it more bluntly: the choice is between keeping your models as IP or leaking them, and if you leak it, it's a one-way door. You're done, in some sense.
When the person selling you the tool and the person criticizing the tool agree on the risk, that's probably more than just a little spin.
What actually leaks when your team works inside an AI tool
Here's the part nobody puts in the onboarding deck: the leak almost never looks like a leak.
It looks like a marketer attaching a client's unreleased positioning into a chat window to get some help tightening the language.
It looks like an analyst dropping a competitive research summary into a general purpose tool because the free version’s right there and the approved one requires a ticket.
It looks like someone pasting three paragraphs of proprietary methodology into a tool to "just clean up the tone”—the same way a recipe gets handed to a friend who hands it to another friend, until it's on a menu three towns over with no attribution and no royalty and no single moment anyone can point to as the theft.
Every one of those hand-offs felt harmless in the moment. But that’s exactly the mechanism. Nobody sits down and decides to give away the firm's edge. They decide, fifty times a day, to save four minutes.
The consumer version of a general purpose AI tool has terms of service most employees have never read, and even fewer moments of pause before pasting.
What goes in doesn't have to be stolen to be gone. It just has to be somewhere you no longer control.
Nobody sits down and decides to give away the firm's edge. They decide, fifty times a day, to save four minutes.
AI vendor selection is a fiduciary decision, not an IT ticket
Most firms still route "can we use this AI tool" through the same lane as "can we use this project management app." It’s the entirely wrong lane. A project management app doesn't ingest your unreleased research, your client's confidential positioning, or the language you haven't cleared with Legal yet. A general-purpose AI tool, used carelessly, can do all three before lunch.
That makes the vendor decision a governance call, not a procurement afterthought—the same category as choosing who holds client data or who audits the trade blotter. The questions that matter aren't "is it fast" and "do people like it." They're: what happens to the data after it's submitted, who can see it, is it used to train anything, and what happens if you need to prove, later, that it never left your walls.
One firm's compliance team recently discovered its content group had been using three different consumer AI tools for eight months, none of them vetted, none of them logged. Nobody had done anything malicious. Nobody had done anything anyone would have approved, either, if anyone had asked.
The uncomfortable truth is that most vendor-risk processes were built for a world where the risky software at least announced itself. A new trading platform gets a security review because everyone knows it's touching sensitive data. A chat window that looks like a search bar doesn't trigger the same instinct, even when what's typed into it is worth more than what's flowing through the platform next to it. The interface is calm. The exposure isn't.
A governance checklist for what data can touch which tools
Not every piece of content carries the same risk, so not every tool needs the same clearance. A simple tiering makes the rest of this tractable:
- Public-facing, already-published material. Low sensitivity. Fine for most general AI tools—repurposing, summarizing, translating.
- Draft copy and internal talking points not yet cleared. Medium sensitivity. Approved enterprise tools only, with data-retention terms your legal team has actually read.
- Client-specific research, proprietary positioning, unreleased methodology. High sensitivity. Enterprise tools with contractual no-training guarantees, or no AI tool at all until one clears review.
- Anything covered by an NDA, a regulatory filing, or a client confidentiality clause. No general AI tool, full stop, regardless of how convenient it looks at 4:45 p.m. on a Friday.
Four tiers, one page, postable next to the espresso machine. The point isn't the document. It's that everyone can name which tier they're working in before they paste anything.
Restricting AI tools isn't the fix
Here's where most firms stop reading and start drafting a policy: ban the tools, block the domains, done. It feels responsible. It's also the wrong move.
A ban you can't enforce doesn't stop the leak. It just stops you from knowing where it is. Employees who've already built a habit of pasting into a convenient tool don't stop pasting when you tell them not to—they just stop telling you about it. The behavior goes underground, the logging disappears entirely, and the firm that thought it solved the problem now has less visibility than before the policy existed.
The smarter move is amnesty first, restriction second. Ask people, without penalty, what they've actually been using this month and for what. You'll get an honest answer exactly once—while it's still safe to give one. Build the checklist and the approved-tool list off what's actually happening in your building, not off what the policy memo assumed was happening. Govern from the truth, not from the fiction that a ban already fixed it.
Where to start Monday
Start with the question nobody's asked out loud yet: what AI tools has your content team actually used this month, and on what.
Ask without penalty, this week, before you draft a single policy line. Then map what data touched each tool against the four-tier checklist above—you'll likely find most of it is fine, and a smaller slice needs to move to an approved tool immediately.
Only after that mapping is honest should the policy get written. A governance framework built on what people admit beats one built on what a memo assumed, every time—the same discipline behind publishing faster without adding risk applies here too.
That's the fix we're not selling. That's just where to look first.