Employees reported pasting financial information into AI tools at nearly 3X the rate security teams estimated: 25.6% of knowledge workers admitted to doing it, while only 8.9% of security leaders estimated that financial data was leaking into AI apps.
Picture it: a financial analyst is under a deadline to deliver specific analysis to the CFO. He copies a table from a shared spreadsheet, pastes it into a common AI tool like Gemini, gets an output, and moves on. The same file being downloaded or sent to a Gmail address might trip a DLP rule. But financial data leakage via a copy and paste into a browser tab? That doesn't register anywhere.
That's what Neon Cyber found in their surveys of knowledge workers and IT and security leaders. Overall, workers reported far more sensitive data moving into AI tools than security teams estimated, and the categories that were underestimated track almost exactly with how well-instrumented each data category already is.
What kind of data is going into AI tools?
Sensitive company data, from source code to financial information, leaked into AI tools at much higher rates than security estimated.
- Financial information: 25.6% of workers reported pasting or uploading it; security estimated 8.9%.
- Sensitive or regulated data: 20.3% reported vs. 14.2% estimated.
- Source code and scripts: 13.2% reported vs. 8.3% estimated.
- Customer or prospect data: 25.1% reported vs. 23.1% estimated.
- HR or employee data: 20.3% reported vs. 18.3% estimated.
- Credentials and API keys: 11.0% reported vs. 11.8% estimated — the one data category where security's estimate actually edged slightly above what workers admitted to.
If your enterprise is still thinking that you’re able to see and stop data leakage into AI tools, you’re already at risk. Depending on what's in that data, regulatory exposure may be part of what's at stake — and beyond that, the data now lives in AI tools you don't manage, subject to whatever data-handling terms that tool has, which most employees never read before pasting.
The one data category where security and knowledge workers seemed to align was around credentials and API keys. Why? Because these are issued, rotated, revoked, and logged as routine IT and security operations. When security estimates credential exposure, they're drawing on a data source that they have direct line of sight to. That's why the estimate lines up with what workers actually report.
Financial figures, source code, and customer records live somewhere without that history: inside SaaS applications, spreadsheets, code repositories, and CRMs that were never built to log a copy-and-paste event. A number moving out of a spreadsheet and into a prompt window leaves no record anywhere security currently looks.
The paste event your DLP never logged
Closing this requires visibility positioned at the point where the copy-and-paste actually happens — inside the browser session, independent of whether the source system was ever built to log that action. A DLP rule extended to more file types doesn't help here, because the risk was never about file transfers. It's content moving out of an application and into a prompt, in real time, in a place none of the existing instrumentation was built to watch.
Get the full reports
Dive into the full results by requesting your copies of each report:
Knowledge Workers Research Report: Quantifying Shadow AI Risk in the Browser
Security & IT Leaders Research Report: Concern to Control - The Shadow AI Blind Spot
*Methodology note: this comparison sits alongside Neon Cyber's companion workforce survey (n=227) and this report's security/IT leader survey (n=169). The two samples aren't population-matched, so the figures referenced are directional, not an exact head-to-head measurement.
FAQs
How can security teams get an accurate picture of what data is going into AI tools?
The most reliable approach is visibility positioned at the point where the data actually moves — the browser session — rather than relying on estimates extrapolated from adjacent, differently-scoped controls like file-transfer DLP or email security. Direct instrumentation at that point removes the need to estimate at all.
How reliable are workforce self-reports of AI data leakage compared to security estimates?
Self-reports carry their own limitations — underreporting is possible for sensitive categories — but they're grounded in direct knowledge of the action taken, while security estimates for uninstrumented categories are inferred rather than measured. Neon Cyber's report treats the comparison as directional rather than exact, given the two survey populations aren't identically matched.
Why is source code harder to estimate accurately than customer data?
Source code usually lives in repositories and IDEs with no existing logging tied to copy-and-paste actions into a browser, while customer data often sits in a CRM with some access logging already in place. Neon Cyber's research found source code exposure underestimated by security at nearly double (13.2% reported vs. 8.3% estimated), while customer data estimates were much closer (25.1% vs. 23.1%).