October 1, 2026
Operational Security in DeFi: Why Most Losses Are Self-Inflicted
The cryptography almost never fails. The people operating it do.

By Cipher Chain Capital
12 min read
- 1 Introduction: The Wrong Mental Model
- 2 Part I: Start With the Threat Model, Not the Tools
- 3 Part II: Key Generation Is the First Irreversible Decision
- 4 Part III: Wallet Architecture Is About Surviving Failure, Not Just Preventing Theft
- 5 Part IV: Storage Is Where Physical Reality Reenters the Picture
Introduction: The Wrong Mental Model
Ask someone how money gets stolen in crypto and they will describe a hack. A genius in a dark room breaking the code, cracking the wallet, beating the math. It is a comforting story, because it makes the loss sound like lightning. Unavoidable. Nobody's fault.
The record says otherwise. In 2025, more than $3.4 billion in crypto was stolen, according to Chainalysis, and the largest single loss was not caused by broken cryptography. In February, Bybit lost roughly $1.5 billion after signers approved a transaction they could not properly verify through an interface that looked exactly as it always had. The lesson was not that the math failed. It was that the operation around the math failed.
Bad key generation. Insecure signing environments. Social engineering. Lost backups. In most catastrophic losses, nobody broke the cryptography. The system worked exactly as designed. The humans did not.
The core idea: operational security is something you design and operate. Nobody sells it in a box. It is people, processes, environments, and failure assumptions, and if any one is weak, the whole structure is weak. The good news is that operational risk is one of the few risks teams can meaningfully reduce with discipline. Security does not require paranoia. It requires repeatable habits.
One note on the lens. This article is about how people attack you and your firm. The risk buried inside the protocols you invest in is a different problem, and one we have written about before. The examples lean DeFi, but the same mistakes drain DeFi treasuries, exchanges, and custody operations alike, which is part of the point. It all rolls up to the question underneath everything we write about risk: "Where can my money actually be lost?"
That makes operational security less mysterious than it sounds. The goal is not to predict every attack. The goal is to build a system that can survive ordinary mistakes, rushed decisions, and partial failures before they become catastrophic.
Part I: Start With the Threat Model, Not the Tools
Every serious security system begins with a threat model: who can attack you, what they can realistically do, what failure would look like, and what you are actually protecting. Skip that step and every other decision is a guess.
The profiles are wildly different. A solo user worries about phishing links, drainers, and a seed phrase caught on camera. A fund operator worries about targeted social engineering and every provider that touches the money. A DAO worries about public signers pressured one by one. An automated strategy worries about hot keys on someone else's server.
Copying someone else's setup without copying their threat model is security theater. The solo user who builds a five-of-seven multisig has built a system he will quietly bypass within a month, because it does not fit his actual life.
Without a threat model, you protect the wrong things very well and the important things not at all: the trader with an elaborate hardware wallet ritual who grants unlimited approvals to every new farm.
Part II: Key Generation Is the First Irreversible Decision
Most people treat key generation as setup. It is not. It is the birth of the entire security model. Every signature, backup, rotation plan, and recovery path traces back to that first moment.
Two things matter: entropy and environment. Entropy is the randomness the key comes from. Weak randomness creates a key that is more vulnerable to brute force, and the danger is that weakness leaves no visible scar. Environment is everything around the key while it exists in plaintext. If a seed phrase is generated on a daily-use laptop, photographed, copied into a notes app, or briefly synced to the cloud, the damage may already be done. If the machine was compromised when the key was created, everything built on top of that key is compromised too. There is no patch for a key born in a dirty room.
"I'll rotate later" rarely fixes the real problem, because rotation often happens on the same laptop, through the same browser, with the same habits. A new key can inherit the same infected process.
The fix is unglamorous: use a clean or dedicated device, stay offline where possible, keep cameras out of the room, and never let seed material touch cloud storage, screenshots, messaging apps, or general-purpose browsers. No shortcuts, because the shortcut is the vulnerability. Keep a record of every key from the moment it exists: what it controls, where it is stored, and who can access it. You will need that list later, usually in a hurry.
Part III: Wallet Architecture Is About Surviving Failure, Not Just Preventing Theft
Most setups are designed to stop attackers. The better goal is surviving mistakes, accidents, and partial compromises too.
A single key wallet is one secret: one phishing click, one bad backup, one house fire, and the money is gone. A multisig spreads authority across several keys and requires a quorum to move funds. The quorum math matters. One-of-one has a single point of theft. Three-of-three has three single points of loss: misplace any key and the funds freeze forever. Two-of-three survives either a stolen key or a lost key, though not both at once.
Put the keys in different failure domains. Two keys in the same house share a fire. Two signers on the same flight share a crash. Separation protects against floods, seizures, coercion, and plain bad luck, all far more common than movie-grade hackers.
The same redundancy logic applies to service providers. If your wallet infrastructure provider is compromised, down, or locking you out, you need a path to your funds that does not run through them. A second provider already stood up is the ideal. At minimum, structure internal backups so you can sweep funds to a clean destination without the provider's stack.
Part IV: Storage Is Where Physical Reality Reenters the Picture
Most security planning focuses on remote attackers and skips the physical layer. Fire, water, and plain misplacement destroy backups just as reliably as thieves do.
Each storage method has a specific weakness. Paper does not survive a house fire or a burst pipe. Steel plates do, and they also sit unencrypted in a drawer, so whoever finds one controls the funds. Encrypted drives fix that and add a new secret: the password now needs its own backup, and drives fail. Home safes can be forced or carried out. Bank deposit boxes keep branch hours and can be frozen in a dispute.
Storage also degrades on its own. Thermal paper can fade to blank within a few years, ink runs, cheap metal corrodes, basements grow mold. Most people store a backup once and never check it again, so the condition of a five-year-old backup is a guess.
Tamper evidence has one practical job. A sealed bag does not stop a thief. It tells you whether someone has already seen the contents, which lets you rotate the key before anything moves rather than after. That early warning is cheap to buy and worth a great deal.
Vitalik Buterin's public setup is a useful illustration of where this ends up. He keeps the large majority of his funds in a multisig, with some keys held by him and the rest spread across people he trusts, and he has said not to reveal who those people are, even to each other. The design starts from the assumption that he himself is the most likely point of failure, and removes it. It is less convenient than one steel plate in one drawer, and that is the trade. A single backup is a single thing to find, lose, or burn. Distribute it across people and places and no one location and no one person can compromise or rebuild the whole.
Part V: The Signing Environment Is the Real Attack Surface
Storage gets the attention. Usage is where the money actually leaves.
Most theft does not happen because someone guessed a private key. It happens because someone signed a bad transaction in a compromised environment. Browser wallets that render whatever a malicious page feeds them. Clipboard malware that swaps addresses mid-paste. Transaction substitution, where the screen shows one payload and the wire carries another. Blind signing, where the hardware wallet displays a hash no human can read. Social engineering around all of it.
Map the surface in three layers: your own devices and operations, those of every provider you rely on, and the smart contracts both touch. Take a Safe multisig. You are exposed to how you interact with Safe, then to how Safe runs its own interface infrastructure, then to the contracts underneath.
Bybit ran straight through the second layer. Forensic reviews by Sygnia and Verichains found that attackers compromised a Safe developer's machine, injected malicious JavaScript into the Safe interface, and served the payload specifically to Bybit's signers. The screens showed a routine internal transfer. The hardware wallets required blind signing, so the one device that could have caught the lie displayed unreadable hex instead. Every signer approved. That one transaction quietly handed the attackers control of the wallet, and roughly $1.5 billion was drained shortly after. Bybit's own systems were never breached. No key was stolen. A centralized exchange, drained through DeFi-native tooling, by an attack on a vendor.
The defense is isolation plus verification, and the encouraging part is that it is entirely procedural. Dedicated signing machines that do nothing else, airgapped workflows for large movements, decoding the calldata, confirming hashes and destinations out of band. None of it requires new technology. A team that adds these steps closes the exact gap that cost Bybit a billion and a half. A perfect vault does not matter if the door is open every time you use it, and the door is something you can learn to check.
Timelocks close the same gap from the other side. Instead of trying to make every signature perfect, a timelock delays execution after approval, so a large movement sits in a public queue before it can settle. A signer who blind-signs the wrong payload then has a window where the transaction is visible, decoded, and cancelable. A timelock would not have stopped Bybit's signers from approving, but it would have held the change in a public queue, giving them time to see what they had actually signed and cancel it before the attackers ever gained control. The cost is speed, which is the trade running under all of this. Every control that buys you a chance to react also slows you down when nothing is wrong.
Part VI: Social Engineering Attacks the Process, Not the Wallet
Security systems are operated by humans, and humans are good at their jobs in ways attackers exploit. Responsiveness. Trust. Deference. Speed under pressure. Social engineering weaponizes the traits most teams hire for.
The patterns repeat. Impersonation: the "CEO" on Telegram who needs a transfer approved before his flight. Urgency: act in the next ten minutes or the position gets liquidated. Authority: compliance needs access, the auditor needs the export, the investor needs the wallet list. None of these attacks begin by breaking the wallet. They begin by bending the process around the person holding the door.
This is why "people are the weakest link" is too shallow. The real problem is not that people are careless. It is that many systems quietly depend on one person being perfect at the worst possible moment. One tired signer. One rushed approval. One employee who does not want to be the reason a deal, withdrawal, or client request gets delayed.
North Korean operatives have repeatedly gotten hired as remote IT workers inside crypto companies, exchanges, custodians, and web3 firms, then used that access to stage thefts from the inside. It is often easier to enter through hiring, trust, and routine access than to attack the cryptography directly.
The institutional answer is to assume the mistake will happen and design around it. Role separation, so no one person can move funds alone. Approval rules that do not change because someone sounds urgent. Clean escalation paths, so a junior employee can stop a suspicious request without feeling like they are challenging authority.
All of this is done to protect the team. Good process takes the weight off any single person and turns a human weak point into a shared, defensible system.
Part VII: If You Are Not Watching, You Are Not Defending
You cannot respond to what you cannot see. That is the whole concept of monitoring, and most operations skip it. Logs are the record of what happened and monitoring is the system that tells you when the record starts looking wrong.
Onchain, the events worth alerting on are knowable in advance: a new approval from a treasury address, a change to a signer set, a timelock queue entry, outflows above a threshold. Even dumb alerts beat silence. The goal is knowing what normal looks like so abnormal is loud.
Most post-mortems read the same way: "We discovered the issue several hours later." Sometimes days. Meanwhile the funds cleared a bridge in minutes.
Detection speed is often the entire difference between a scare and a catastrophe, and it is one of the highest-leverage things you can buy. An attacker noticed at the approval stage gets nothing. Good monitoring turns most incidents into footnotes.
Part VIII: Assume Compromise and Plan for It
Good security assumes failure. It is the part most operations skip, and probably the most important one here.
A real compromise plan has parts you can point to. The key inventory from Part II, so when something feels wrong you know exactly what could be exposed. Named roles, so everyone knows who freezes what and who calls whom. Communication channels that do not run through possibly compromised systems, because if your Slack is the breach, your Slack cannot be the war room. And rehearsals, because a plan that has never been run is a document, and documents do not respond to incidents at 3am.
The freeze itself deserves its own rule. Make pausing easy and unpausing hard. The ability to halt the protocol should sit behind a low threshold, a single trusted signer or one-of-five, because an emergency at 4am will not wait for three or four people to wake up and coordinate. Pausing moves no funds, so the low bar costs almost nothing. Unpausing is where the stakes return, so route that through a higher admin quorum. A pause that needs the same signatures as a transfer is close to useless at the one moment it exists for.
The scenarios worth rehearsing are mundane, which is what makes them likely. A signer stops answering. A device gets seized at a border crossing. A backup comes out of storage with a broken seal. Something feels off, you cannot prove anything, but you know something is wrong.
A plan for these scenarios is not hard to build. It is mostly a matter of writing down who does what and running it once before you need it. Teams that do this handle a bad day calmly. Teams that say "we'll figure it out" are improvising while the clock runs, which is the one variable you can remove in advance.
Part IX: The Forgotten Risk of Old Devices and Old Data
Deleting a file does not remove it. Deletion only removes the reference, while the data sits on the disk until something overwrites it. Forensics recovers it routinely.
Now count the places your security model has lived. Old laptops. Retired phones. Cloud snapshots nobody remembers taking. The backup-of-a-backup in a drawer. Each is a frozen copy of your operation, possibly including key material, a photographed seed, an export made just in case. Yesterday's hardware is tomorrow's breach.
In practice, the cleanest answer is often to stop trusting the old keys entirely and migrate funds to a fresh setup, rather than betting the treasury on a wipe you cannot verify. Anyone who has done this knows it is a slog, especially once service providers are involved, since whitelists, integrations, and counterparty addresses all move with you. Annoying, genuinely. Still cheaper than the alternative.
Operational security covers how things end, and most things end in a drawer.
Part X: Audits Are Not About Trust, They Are About Blind Spots
Every serious security domain builds in third-party review. Aviation, banking, nuclear power. The reason is the same everywhere: people are too close to their own systems to see what they missed.
Crypto has internalized exactly one kind of review, the code audit, and it covers the smallest slice of the problem. An infrastructure audit examines the machines and cloud accounts around the contracts. A process audit examines how humans actually handle keys, approvals, and recovery. Nearly every loss in this article lives in those last two, and most operations have never had either reviewed.
Frameworks exist. The CryptoCurrency Security Standard benchmarks these controls from key generation through decommissioning, and there is a growing industry push to standardize operational due diligence for DeFi, because diligence today mostly stops at the contract audit.
The value also runs outward. A review tells you what you missed, and it tells everyone else something too. Conforming to a standard like CCSS signals to depositors and users that the team is aware of operational risk and was following these practices at the time it was checked. That is a snapshot, not a promise about next week, and it should be read as one. But in a space where most teams have never had their process or infrastructure looked at by anyone, being able to point to a review is a real signal.
An audit will not prove you are perfect. It will find the thing you did not know you were wrong about. That is the entire value.
Conclusion: Crypto Security Is a System, Not a Product
You do not buy security. You design it, operate it, maintain it, and rehearse its failure modes. The wallet, the provider, the intern with Telegram access: all components of one system, and the system is only as honest as its weakest assumption. The encouraging part is that every one of those components is something you can choose and improve.
Look through the catastrophic losses of the past few years and genuine brilliance is rare. What you find instead is ordinary operational mistakes inside systems that were never designed to survive them. A key generated on the wrong machine. A signer who trusted a screen. A plan that was never written down. That is the good news hiding in every post-mortem: these were preventable, which means yours are too. The work is unglamorous and the payoff is quiet, and quiet is exactly what you want.
That is why the lesson is not despair. It is ownership. The same operational gaps that make losses possible are the places where better habits, better process, and better review can actually change the outcome. In DeFi, the cryptography is usually the strongest part of the system. Everything around it is where the money is lost, and everything around it is where you get to build.
Author: Austin Liu | Telegram | Instagram | Medium | LinkedIn | X