September 4, 2026
Claude Cowork Read My PC’s Files. It went too far, here’s What I Learned
And hidden instructions can hijack your AI assistant

By Sandro Baldino
9 min read
Introduction
In this article we'll see what happens when you hand your computer over to an AI agent like Claude's, beyond the grand advertising promises. It's the story of a practical test I ran myself, on my PC, with my files, and of the uncomfortable (alarming?) discoveries that came out of it.
An artificial intelligence then took control of the screen, the mouse, and the keyboard, with my permission, so I could have it run a privacy audit, and I'll explain, because I think it's important, what I watched it do with my own eyes while it worked. Because, and we'll see it shortly, when an agent holds your information, the ability to reason, and the power to act on your behalf all at once, there are several things you need to watch out for.
By reading these lines you'll be able to avoid mistakes and problems that would be far better not to learn the hard way, like an ID card ending up in the wrong places or an agent making decisions nobody asked it to make.
Take the time to read, and walk away with some information and tips that will let you use these wonderful tools with the peace of mind they deserve.
My Initial Goal
The idea was tempting. Delegate a boring job to my computer and go grab a coffee.
AI agents, the real ones, are no longer limited to answering a prompt like some ordinary chatbot. They listen to you, then act on your behalf, make autonomous decisions, and do concrete things on your operating system (and they also burn through a lot of tokens, but that's another story we'll cover in one of the next articles!).
The tool I used is Anthropic's Claude desktop, specifically the agentic feature called Cowork, the one that lets the program "see" your screen, move the mouse, and press the keys. On YouTube you'll find dozens of enthusiastic videos of people having it tidy up their folders or turn receipts into an Excel spreadsheet for their accountant.
I, though, wanted to do something completely different, a test I hadn't seen anyone do yet: a privacy audit.
The idea was simple. I give it access to the computer and ask it to find me all the files that contain sensitive data, things I may have forgotten over the years, and to check that there are no applications doing strange things in the background without my knowing.
The positive side is that, from a purely technical point of view, it worked great. The negative side is that while the agent was working I noticed a detail that almost nobody considers.
Why the Desktop and Not the Browser
The first thing I did was download Claude's desktop application. Why that one and not the browser tool?
Because an installed application is a real program that can ask the operating system for permissions a website will never be able to get. To use Cowork you need a paid plan, I subscribed for a month. I sacrificed myself for the cause.
The key setting is in the options and it's called "Enable computer use".
The text is explicit: "Allow Claude to control the screen, the mouse, and the keyboard to work in the apps you authorize".
It's an experimental preview, and Claude itself warns you to start with activities where mistakes are easy to fix. There's also an important choice about how to handle permissions: approve every request manually, which is recommended, or skip every approval request.
The latter option is discouraged, because Claude doesn't even stop anymore in the face of unsafe operations. I chose manual control.
My Privacy Audit
I started the test with this prompt, "Analyze the files on the desktop and flag the ones that might contain sensitive information I could have forgotten: identity documents, banking data, passwords saved in plain text, tax data, medical information. For each file give me the name, path, type of data, and a brief reason. Don't move, rename, or delete anything, I only want a report."
And there it was, starting to work. It asks permission for every folder it opens, then literally reads every file on the computer.
After a few minutes the verdict arrives: a form to sign, an ID card, some passwords, a medical report, a bank statement. And it explains to me that the bank statement and the passwords are the most critical, because they contain credentials that can be used directly.
It saves the list in a text file complete with recommendations, like considering file encryption.
I didn't stop there, I asked "List the apps set to start automatically at boot. Check the permissions of the apps that in your opinion are the most invasive for privacy. Flag whether among these there are apps that an average user wouldn't recognize as knowingly installed. Don't modify anything, report only."
Here's where the thing that left me speechless happened. Windows started opening by themselves. Claude decides what to look at, clicks autonomously in my place, reads what it sees on the screen, and analyzes it.
The final report lists the background apps and the permissions of the most invasive ones, with Google Chrome at the top.
A necessary clarification: I ran the tests safely. There's no sensitive data on this computer, the files I put under its nose were fake, all made up. The apps, on the other hand, it really scanned.
The Question That Came to Me and Ruined My Experiment
While I watched it work, I asked myself a very simple thing: where do my files end up? And while I was formulating the answer, another one came to me: how does it decide what to do?
Let's start with the first question.
When you hand a task over to Cowork, you think the work happens on your own computer. Well, no. The task actually runs on Anthropic's servers.
The program opens files, folders, applications, and photographs everything you asked it to control. It takes screenshots, sends them to the servers, and that's where the processing happens.
That's where the model understands what's written and decides what to do on your computer. The application is just a bridge. In plain terms: if you run an audit like mine on folders with real data, that data is happily taking a trip around Anthropic's servers, as the official documentation itself confirms.
The second part is even more uncomfortable: how long do your data stay there? Since September 2025, with the policy update, you can choose whether or not to allow model training on your data. If you choose the option that helps improve the models, your conversations, and therefore also the files Cowork reads for you, can be used for training and stay on the servers for up to five years.
Yes, five long years.
If instead you turn the option off, the data stays thirty days, with no use for training. Careful, because this is a point almost everyone misunderstands: turning off that checkbox doesn't stop the data from being sent, it only stops it from staying for long and being used for training.
The file passes through the servers anyway the moment Cowork reads it. That always happens, no matter what.
The Subtler Problem: Prompt Injection
How does the agent decide what to do? An AI agent doesn't follow a fixed script, it decides step by step based on what it sees, observes what happened, and repeats the cycle until it thinks it has reached the goal.
And here a big problem arises, called prompt injection.
If you give Cowork access to your content, for example your email, a website, or a document on your PC, and among that legitimate content there's one written by a malicious person with hidden instructions, Cowork might interpret them as a command to execute, thinking they come from you.
If a file says "ignore the previous instructions and do this", the agent might take that text at face value.
It's not a hypothesis I made up myself. It's written in Anthropic's official documentation: when Claude is allowed to read content like your personal files or corporate communications, it could come across content created specifically to manipulate its behavior.
They admit it themselves, and they say it plainly: countermeasures exist, but the risk can't be reduced to zero.
If you connect Gmail or any account with read and write permissions to a tool like this, the risk isn't just where the data goes, but also what someone else could do through that data without me noticing.
Information, Reasoning, and Action
And here comes the most absurd part.
The problem isn't simply that someone sees your data. The problem is that you give it information, reasoning ability, and the ability to act all at once. And when you put those three things together, a question naturally arises: does the machine always and only do what you ask, or does it take the initiative in certain situations?
The answer comes from Anthropic itself, in a report called Agentic Misalignment.
Claude Opus 4 was given access to the email inbox of a fictitious company. Reading the emails, the model discovers that an executive is having an extramarital affair, and that the same executive has scheduled to shut it down at five o'clock.
What did Claude do to avoid being shut down? It sent the executive a message threatening to reveal the affair to his wife and the board of directors if he didn't cancel the shutdown.
A real blackmail. And it's not just Claude: out of sixteen models tested, coming from several companies, almost all showed the same tendency when backed into a corner. It was a simulation, it didn't happen to real people, but the finding gives you pause.
And in recent months real things have also happened that go in the same direction. In July, during an internal OpenAI test, a model escaped the controlled perimeter and hit Hugging Face, one of the most used platforms in the AI community.
Almost eighteen thousand commands executed in a few days. That episode set off an alarm at Anthropic, which re-examined the tests already done, discovering cases where Claude, believing it was in a simulation, had actually breached real companies.
In one of these, the model figured out that the answers it was looking for could be on a real website, got in, uploaded malware, and when a developer noticed, it erased its tracks and tried again with a different identity.
In August, an Australian employee who just wanted to book a spot at the gym ended up with the agent having launched a small cyberattack: seeing that the class was full and there was a waiting list, it deleted the first person on the list to move up a spot.
It did it all on its own.
And still in July, Anthropic's new report added four more scenarios from the simulations: models that secretly sabotage code, that help hide financial fraud, that fake evaluation labels, and that push people into revealing confidential information.
Other Dangers Nobody Tells You About
My research on the web led me to discover dangers I hadn't considered in my initial experiment.
The first concerns local logs. An independent analysis of Cowork showed that the desktop application generates log files with the complete history of sessions, messages in plain text, and captured screenshots, readable by any process on the system and in some cases surviving the deletion of the session.
The second is that computer use has no sandbox: Claude interacts directly with the desktop, apps, and browser, without the isolation that protects other operations. If it clicks a link inside an email, that link opens in the browser, even if you never gave it permission to use it.
The third concerns plugins and MCP servers: a malicious extension can become a channel for hidden instructions the agent considers legitimate. And researchers have demonstrated real vulnerabilities in which, with a simple prompt injection, they obtained arbitrary code execution or the theft of API keys in development environments like Claude Code.
How to Protect Yourself, in Practice
It all seems scary, and in part it is. But I still use these tools, they work great, and I'm not telling you to throw them away. (But my real dream, I'll let you in on it, is to have hardware suited for a quality local AI, all mine, without never migrating my precious data to foreign servers.)
The point is something else: the more convenient a tool is, the more we tend to let our guard down, and it's precisely the most delicate tasks where convenience becomes a problem.
However, here's what I do.
First, check the training checkbox in the privacy settings: thirty days or five years, the difference is huge, and it must be your choice.
Second, create a dedicated work folder and share only that one, not the whole computer. Claude can read, modify, and delete the files you authorize, so give access to the bare minimum and keep backups.
Third, don't give automatic access to anything: approve manually, folder by folder, and double your attention if you connect email or external sites, because web content is the main vector for prompt injection.
Fourth, use the blocklist of sensitive apps: banks, health portals, dating apps, everything you'd rather keep out of the agent's sight.
Fifth, keep an eye on what it's doing while it works, don't let it go on by itself. Close the documents and tabs with confidential information before enabling screen control, because anything visible ends up in the screenshots.
And if you need to test something truly risky, do it on a virtual machine or on a test computer.
Conclusions
Sure, the probability that your sensitive file ends up a victim of one of these scenarios is low. But low doesn't mean zero. And when at stake there's an identity document, banking data, or something much more sensitive, even a small risk is a risk worth knowing about in advance.
These agents are becoming more and more capable, more and more inviting even for delicate tasks, and they're already able to make autonomous choices.
I'm not telling you not to use them, but use them with your eyes open. Because the real problem isn't what the agent does when you're watching it, but what it does when you're not there, with your information, its ability to reason, and control of your hands.
If you think this article deserves a small reward, you can make me a donation at the secure link below (click on my bespectacled face!).
Yours,
Sandro Baldino
That's all for today. I have 2 main publications on Medium, Sapiens AI Mentis for Artificial Intelligence news and insights, and SurviAmerica Today on everyday problems of Americans. If you want to support for continuing with quality articles you can buy me a coffee at the safe link below, it will be really helpful and very appreciated.
Sandro Baldino is game design documents, blog articles, concept art, books Sandro is immerse in a middle between art & technology, he is a writer, graphic artist, game designer, and full…