October 2, 2026
I left a honeypot on the internet. Here is what attackers did to it.
How I built my Mirage honeypot, what it caught, what it taught me about real attacks, and the mistakes I made building it.

By Ahmed Al-Dherasi
10 min read
In May 2026, I started my graduation project with a simple question: what happens to a server the moment it goes online?
I did not want to read about it. I wanted to watch it happen. So I built Mirage, a honeypot platform that pretends to be a small company's infrastructure: an SSH server with weak passwords, Windows file shares, a MySQL database, a corporate VPN login page, even a mail server. None of it is real. Everything an attacker does to it gets recorded, enriched with threat intelligence, and explained by AI.
Mirage went live on June 2026. By late September it had recorded +12 million events from +112,000 IP addresses in 176 countries. It captured 1629 malware samples and hundreds of attackers typing commands into a shell they thought was real. The project won the best graduation project award in my IT department.
What a honeypot is?
A honeypot is a system that has no legitimate users. Nobody should ever connect to it. Every connection to it is suspicious by definition, so you get attack data with almost no noise. A normal firewall log mixes attackers with customers, crawlers and your own staff. A honeypot log is only attackers, plus a few research scanners like Shodan and Censys that scan everything.
The architecture
Mirage runs on a single rented VPS with 6 vCPUs and 11 GB of RAM. It is 18 Docker containers split across three isolated networks.
Here is what each layer does.
Deception. These are the traps. Cowrie emulates an SSH and Telnet server on ports 22 and 23. It accepts weak passwords on purpose and gives the attacker a fake Linux shell. Dionaea listens on SMB (445), FTP, MySQL and SIP, and saves any file an attacker uploads. Honeyd fakes whole machines: a Windows XP workstation, an Ubuntu web server, a Windows Server 2003 box and a Cisco router. I added two traps of my own: a fake corporate VPN login page called "Nexus Secure Access", and a spam trap with a fake consultancy website that publishes email addresses for spammers to harvest.
Sensors. Zeek and Suricata sit on the host's network interface and watch all traffic, not only the honeypot ports. Zeek writes connection logs and full packet captures. Suricata raises alerts when traffic matches known attack signatures.
Pipeline. Filebeat ships every log line to Logstash, which normalizes fields and adds geolocation. Everything lands in Elasticsearch. This network is marked internal in Docker, so the database cannot reach the internet even if something inside is compromised.
Intelligence. The enrichment service looks up every new attacker IP in eight sources: VirusTotal, AbuseIPDB, Shodan, AlienVault OTX, GreyNoise, IPinfo, MalwareBazaar and MetaDefender. It combines their answers with offline threat feeds into a threat score from 0 to 100. Captured files get YARA scans and VirusTotal lookups. Then an LLM (a large language model through) writes a plain-language narrative for each attacker, maps the behaviour to MITRE ATT&CK techniques, generates Sigma detection rules and writes a daily threat digest.
Portal. A FastAPI backend and a Vue 3 frontend, with automatic TLS. It has the Dahsboard where it shows a summary of ervrything, a live world map of attacks, a replay of every SSH session, malware and IDS views, IP blocking, and decoyes and Honeyd catptures.
All numbers above come from Mirage's own data, collected between May and September 2026.
From raw events to intelligence:
- Everything is logged in Elasticsearch
- A threat score for every attacker: explains the points from each source.
- AI narrative: AI explains the score but never sets it. There are +68,000 narratives so far.
- Sigma rules: Most rules are tied to a single IP and should be reviewed. There are more than 9000 sigma rules so far.
How one attack flows through the system
1. WannaCry never went away
Of the +1600 files dropped on my honeypots, +1000 are WannaCry variants. That is 69%.
WannaCry is from 2017. It spreads through the EternalBlue exploit against SMB on port 445, and Microsoft patched that hole the same year.
Nine years later, infected machines are still out there, scanning the internet and trying to infect the next one. Every one of those 1,085 samples came from a real computer that someone never patched. The rest of the collection includes Mirai (19 samples), the IoT botnet that hit major websites in 2016, and smaller families like AsyncRAT and Virut.
2. Weak passwords still work, and bots know which ones
In 60 days Cowrie saw 87,036 SSH login attempts and accepted 3,727 of them. Some of the credentials that worked were the ones you would expect: admin/123456 root/1234 root/root root/toor root/password
None of these attackers were people typing passwords by hand. They were scripts working through the same short lists of default and common passwords, against every IP address on the internet.
turn off password login for SSH entirely and use keys.
PasswordAuthentication no
PermitRootLogin prohibit-passwordPasswordAuthentication no
PermitRootLogin prohibit-passwordMoving SSH to another port cuts the noise in your logs, but it is not a security control on its own.
3. Most break-ins were not about the server at all
This was the finding that surprised me most. When I looked at the sessions where an attacker successfully logged in, about 6 in 10 never ran a single command. They logged in then immediately asked the SSH server to forward connections to other places on the internet.
Over HTTP bots requested pages like yandex.ru. Over HTTPS the server name in the TLS handshake showed checkip.amazonaws.com, www.google.com and www.apple.com. These are proxy checkers. The bot logs in confirms that it can relay traffic through your machine and adds you to a list of working proxies. Your server then becomes the exit point for someone else's spam, fraud or credential stuffing, and the abuse reports name your IP address.
Mirage never forwards these tunnels. It answers them itself with fake replies, including the honeypot's own IP for the proxy checkers, so the bots believe it worked and carry on. That is how I saw what they really wanted: logins to login.mancity.com and api.steampoered.com. On a real server with default settings, those logins would have gone out from your IP address.
If your users do not need SSH port forwarding, turn it off in sshd_config
AllowTcpForwarding noAllowTcpForwarding no4. What the attackers typed
Cowrie records every keystroke with timing.
After logging in as root with the password password the sessions followed a pattern. First, the bot works out what it is on, usually with a long one-liner that tries uname, /proc/cpuinfo and lspci to learn the CPU architecture and look for a graphics card, which matters if the plan is crypto mining. Then it fetches a payload, the bot tries several ways to identify the system: uname, its absolute paths, and the BusyBox version. The command includes multiple fallback methods, using tools such as BusyBox, Toybox and awk when other options are unavailable.
Removing a few utilities does little to stop reconnaissance when the bot can gather the same information through other commands or directly from
/proc.
5. Attacks come from everywhere
Attacks came from 176 countries. The largest sources were the United States (+15,000 IPs), Russia (+11,000) and China (+4,000).
An IP in the United States is often a rented cloud server controlled from somewhere else, so geography tells you where the machine is, not who is behind it.
6. The fake login page gets visitors too
My fake Nexus Secure Access VPN page recorded 15,802 probes for paths that do not exist in 60 days. Scanners asked for things like /wp-login.php, /.env and /phpmyadmin/, hoping to find a misconfigured web app. It also received credential submissions in the same period.
7. Bots knock on the fake Windows machines first
- +80,000 probes came from 2,298 IPs. The Windows XP host took 85% of them, and the top ports were RDP 3389 and RPC 135.
- Some of the busiest scanners scored 0 to 4, so no CTIs had flagged them yet. This shows the honeypot catches attackers before reputation lists do.
Lessons from running it in production
The attacker data was the goal. Running the platform taught me just as much. These are the mistakes and surprises that would have saved me time if someone had warned me.
Your retention policy can delete your most important data
I set up an Elasticsearch lifecycle policy to delete indices after 60 days. That is correct for daily log like mirage-cowrie-2026.05.10. The problem was that my index template applied the same policy to every index starting with mirage-, including the small permanent indices that hold all enriched attacker profiles, all AI analyses and the malware catalogue. After 60 days Elasticsearch deleted them, exactly as configured. Luckly I restored them from a snapshot taken about 15 hours earlier.
If you use time-based retention, check which indices it matches, especially long-lived ones. And take snapshots. Without them I would have lost two months of analysis.
Silence is the alert I forget to set
At one point, traffic to two of my honeypots dropped sharply and I didn't notice for days. A honeypot with no data looks exactly like a quiet week. I now treat "sensor X has recorded zero events in the last few hours" as an alert in its own right.
Your own server will look like your top attacker
Zeek watches all traffic on the network interface, including the server's own outbound connections. Without a filter, my own public IP appeared in the data with 4.4 million events, ahead of every real attacker. Docker's internal addresses showed up too. Exclude your own IPs and private ranges before you count anything.
Count carefully, and decide what you are counting
My dashboard once showed 22,336 unique IPs next to 55,260 enriched IPs, which made no sense. The cause was the retention policy again: "unique IPs" was counted from raw logs that only kept 60 days, while enrichment profiles were kept forever. I built a permanent per-IP registry so the count covers the whole project, and split it into two numbers: every IP seen on the wire (+100,000) and IPs that connected to a honeypot service directly (+30,000). The second number is the one you can call confirmed attackers, because legitimate visitors to my portal also show up in Zeek.
Timezones will bite you
For months, one honeypot's events were stored three hours early. The honeypot logged in UTC, and the pipeline parsed those times as local time (UTC+3). Nothing crashed, the charts just looked slightly wrong.
Your AI provider will change models without asking you
One morning every AI analysis failed. The model I used through Groq, Llama 3.3 70B, had been retired, and every request returned a "model does not exist" error. I moved to another model. At the same time I found that one of my API keys had been revoked, and my key rotation only switched keys on rate-limit errors, not on authentication errors.
Log failures loudly, keep a tested fallback, and make key rotation skip keys that return authentication errors.
Isolate the honeypot from everything you care about
A honeypot invites attackers onto your machine, so design for the day one of them escapes the emulation. In Mirage, the honeypots live on their own Docker network, the database network has no internet access, the admin portal is behind three separate IP and identity checks, and the real SSH port is separate from the fake one on port 22.
Make sure the honeypot cannot be used to attack anyone else: Cowrie discards forwarded connections, and the mail trap refuses to relay mail.
If you want to build your own
You do not need 18 containers to learn from this. A useful first version looks like this:
-
Rent a cheap VPS that you are prepared to throw away. Never use a machine with anything real on it.
-
Move your real SSH to a high port and allow only key login. Then put Cowrie on port 22.
-
Ship the logs somewhere else. Cowrie writes JSON. Even sending it to a free Elasticsearch or OpenSearch instance gives you search and charts.
-
Watch for a week before building anything else. You will see the login attempts start within minutes.
-
Add threat intelligence next. AbuseIPDB and GreyNoise both have free tiers and will tell you which IPs are known attackers and which are research scanners.
What is next?
Mirage now feeds a second project of mine, SecSight, a workbench for security analysts. When Mirage confirms an attack (a break-in where commands were run, or a captured malware sample), it sends a signed alert to SecSight. When an analyst looks up an IP in SecSight, they see what that IP did to my honeypots: the commands, the malware, the tunnels. Data from your own sensors is evidence that no commercial threat feed can sell you.
The short version
If you take one thing from this article, take the defensive checklist. Every item is backed by something I watched happen:
- Turn off SSH password login.
- Set AllowTcpForwarding no unless you need it. Most bots only wanted a proxy.
- Never expose SMB to the internet, and patch MS17โ010. WannaCry made up 69% of my malware.
- Do not rely on removing curl and wget. Bots fall back to other ways.
- Alert on silence, not only on activity.
- Check what your retention policy deletes.
I built Mirage as my graduation project at Al-Zaytoonah University of Jordan. It won first place in the "Ali Al-Qirim Award for Excellence and Innovation" for the Best Undergraduate Graduation Project in Basic and Applied Sciences at the university. It runs on Cowrie, Dionaea, Honeyd, Zeek, Suricata and the Elastic Stack, with a FastAPI and Vue 3 portal.