The Unaudited Ten Percent

0xIvy
Meme Coins

I remember the first time a probability frightened me more than a certainty. It was 2017, twelve weeks into an audit of 150,000 lines of Solidity, and I had just found a logic flaw that was invisible to every linter we ran. The code compiled. The tests passed. The exploit existed anyway. What scared me was not the bug. It was that the people funding the project had seen a green checkmark and decided, reasonably, to trust it.

I felt that same cold thing again this week, reading Jacob Coxon's post on X. A safety researcher at Anthropic, formerly of OpenAI, wrote that he puts the probability of artificial intelligence "killing all humanity" above ten percent. Hours earlier, another Anthropic employee had announced their resignation, saying the lab was gambling with our lives. Coxon then announced his own departure, calling both Anthropic and OpenAI irresponsible — "racing toward self-evolving superintelligence, using our lives as the stake."

Two people. One number. Zero methodology. And a blockchain column, of all places, is where I want to work through why that last part is the actual story.

The Context Nobody Publishes

Anthropic built its reputation on being the lab that publishes. Constitutional AI. The Model Spec. Responsible Scaling Policies that name capability thresholds and the commitments they trigger. Whatever you think of them, they exist, they are public, and they invite scrutiny.

That brand is why the resignations land so hard. When a safety-first lab loses safety staff who then describe the work as a life wager, the criticism arrives pre-authenticated. The messenger has credentials, equity in the mission, and nothing obvious to gain. OpenAI's name appears in Coxon's post because the two companies are locked in the same race, and everyone in the industry knows it, including the people signing the checks.

But read the coverage carefully and you notice what is missing. No evaluation suite. No red-team appendix. No internal probability model, no reference class, no stated assumptions, no sensitivity analysis, no confidence interval. Not even an operational definition of what "killing all humanity" would mean. Just a number, delivered with the moral weight of a resignation behind it.

I have spent twenty-six years watching code make promises it could not keep, and I have developed a low tolerance for claims that arrive without their receipts. Not because I doubt the researcher's sincerity — I don't — but because sincerity and verifiability are different properties, and only one of them scales.

What a Probability Owes You

A probability is not an opinion wearing a number's clothing. To say ten percent honestly, you need a model of the system, a reference class of comparable systems, priors, and a mechanism for updating. In my Compound governance work in 2020, four of us found a reward-distribution flaw that quietly favored early depositors — the exact opposite of the protocol's egalitarian manifesto. We published the function, the input conditions, the affected block range, and a reproducible test. Anyone could rerun it. Anyone could disagree with our severity rating, and some did, and the conversation got better.

That is the gold standard of a risk claim: falsifiable, reproducible, attributable.

AI extinction risk cannot be tested that way, because the artifact does not exist yet. You cannot run a Monte Carlo on a system you have not built. The honest form of the claim would therefore be a model of the model — assumptions published, priors declared, uncertainty stated as uncertainty. That is not what we got. We got a headline number, and headlines do not carry error bars.

There is a middle ground, and it is the one I have been building toward for the past year. In 2026 I led a six-month open-source effort to put a verifiable AI training corpus on-chain — not the data itself, but its provenance. Dataset hashes. Signed attestations from the curator. Inclusion proofs for every record. A public, append-only log of what changed and when, so that a bias audit performed in March can be checked against the exact corpus that produced the March checkpoint.

I want to be precise about the limits, because my industry oversells this by reflex. Proving where data came from is not the same as proving what a model will do. An on-chain log of training provenance says nothing about emergent behavior at scale, and anyone claiming otherwise is selling you a bridge to a continent that has not been discovered yet. But provenance is a floor, and we are currently standing on dirt. Red-team results could be attested. Evaluation runs could be attested. Capability thresholds could be timelocked, so that a lab's stated policy becomes something you verify rather than something you read. The chain is a notary, not a savior. A notary is still worth having.

There is a deeper problem here, and it is the same wall I hit whenever I price a novel risk. You cannot estimate the probability of civilizational failure from a sample of one. There is no reference class of civilizations that built superintelligence and stopped. So the number is not derived from data; it is a judgment translated into arithmetic. That translation isn't dishonest — actuaries do it too — but it has to be exposed. When I estimate the exploitability of a governance module, I state my assumptions out loud: these roles can be captured, this timelock can be bypassed within N blocks, this oracle is manipulable for X dollars. Strip the assumptions away and my number stops being analysis and becomes incantation.

Where My DeFi Scars Start Talking

Here is the part that should make everyone uncomfortable. Between 2020 and 2022, I watched dozens of protocols report enormous total value locked that existed only because the protocol was paying for it. Yield farming at 400% APY was not adoption; it was a line item. The moment emissions dropped, the "community" evaporated into a Telegram channel full of people asking about the next farm. The metric was real. It was just measuring the subsidy, not the demand.

Safety branding has the same failure mode. A lab can maintain a safety team, publish a constitution, hire a head of alignment, and still be optimizing for the things that actually determine its survival: capability benchmarks and capital. When a safety function's output is unfalsifiable, you cannot tell a genuine constraint on the race from a well-funded marketing department that also writes papers. Safety, when it cannot be measured, becomes a marketing function with a compliance budget.

I watched the same pattern during the modular blockchain summer. Dedicated data availability layers were pitched as non-negotiable infrastructure for every rollup, while most rollups post a few hundred kilobytes a day — within reach of Ethereum calldata plus a blob. The infrastructure was not wrong. It was premature, and prematurity is how you build a cathedral for a congregation that never arrives. Alignment infrastructure carries the same risk: elaborate machinery whose demand is asserted rather than demonstrated.

Which brings me back to the resignations. In the on-chain world, exit is legible. A wallet moves, a validator exits, a delegate withdraws — the act is public, timestamped, provable, and its meaning can be argued about but not its occurrence. A resignation letter has none of those properties. It is a claim with skin in the game and zero verifiability. It can be sincere, strategic, anguished, or all three at once, and we have no epistemically honest way to tell. That is not a reason to dismiss it. It is a reason to weigh the fact pattern rather than the number. Two departures within hours, from a lab whose entire competitive advantage is safety credibility — that is expensive. Expensive signals are worth more than cheap ones. You cannot fake the cost.

The Contrarian Read

Everyone in my feed read these resignations as proof that the labs are reckless. I read them as something less flattering to all of us: the most rhetorically powerful form of safety communication is also the least accountable one. The person who leaves keeps their moral standing and sheds every lever of control. The people who stay inherit the decision, and they stay. If the goal is genuinely to reduce risk inside a lab, walking out the front door is not obviously the action that achieves it.

Second, my own industry should sit this one out rather than congratulating itself. Crypto invented values theater. We priced tokens on governance language we never enforced, subsidized liquidity and called it product-market fit, and shipped "decentralized" systems with a single upgrade key. Code is law only when it aligns with human values — and we of all people should recognize a constitution nobody can test.

Third, and this unsettles me most: an unfalsifiable risk claim does something strange to a debate. It cannot be refuted, so it converts argument into identity. Ten percent becomes a shibboleth. You either accept the number or you are part of the problem. And a number that cannot be checked is a number that cannot be improved.

Takeaway

What would credibility actually look like? Publish the model behind the ten percent. Name the assumptions, the priors, the reference class you borrowed and why. Invite independent auditors with real access rather than a press cycle. Attest the evaluation runs. Timelock the commitments. Let the claim be wrong in a way that can be pointed at and corrected.

We have the primitives. What we do not yet have is the nerve to use them on ourselves.

So here is the question I keep circling at my desk in Denver: if you cannot audit the ten percent, what exactly are you buying — and from whom?

Market Prices

BTC Bitcoin
$75,549.1 -3.91%
ETH Ethereum
$2,396.48 -5.71%
SOL Solana
$96.82 -6.15%
BNB BNB Chain
$712.4 -1.56%
XRP XRP Ledger
$1.28 -11.15%
DOGE Dogecoin
$0.0799 -5.08%
ADA Cardano
$0.1948 -7.24%
AVAX Avalanche
$7.25 -5.08%
DOT Polkadot
$0.9451 -6.35%
LINK Chainlink
$10.88 -6.22%

Fear & Greed

69

Greed

Market Sentiment

7x24h Flash News

More >
{{快讯列表(10)}} {{loop}}
{{快讯时间}}

{{快讯内容}}

{{快讯标签}}
{{/loop}} {{/快讯列表}}

Event Calendar

{{年份}}
28
03
unlock Arbitrum Token Unlock

92 million ARB released

08
04
upgrade Solana Firedancer

Independent validator client goes live on mainnet

12
05
halving BCH Halving

Block reward halving event

15
04
halving Bitcoin Halving

Block reward reduced to 3.125 BTC

18
03
unlock Sui Token Unlock

Team and early investor shares released

10
05
upgrade Ethereum Pectra Upgrade

Raises validator limit and account abstraction

22
03
unlock Optimism Unlock

Circulating supply increases by about 2%

30
04
upgrade Celestia Mainnet Upgrade

Improves data availability sampling efficiency

Tools

All →

Altseason Index

42

Bitcoin Season

BTC Dominance Altseason

Gas Tracker

Ethereum 28 Gwei
BNB Chain 3 Gwei
Polygon 42 Gwei
Arbitrum 0.5 Gwei
Optimism 0.3 Gwei

Market Cap

All →
1
Bitcoin
BTC
$75,549.1
1
Ethereum
ETH
$2,396.48
1
Solana
SOL
$96.82
1
BNB Chain
BNB
$712.4
1
XRP Ledger
XRP
$1.28
1
Dogecoin
DOGE
$0.0799
1
Cardano
ADA
$0.1948
1
Avalanche
AVAX
$7.25
1
Polkadot
DOT
$0.9451
1
Chainlink
LINK
$10.88

🐋 Whale Tracker

🟢
0x0006...5aa4
30m ago
In
4,955.84 BTC
🔵
0x7635...cc92
5m ago
Stake
2,569 BNB
🟢
0xc01c...46bd
6h ago
In
34,697 BNB

💡 Smart Money

0x69e4...f67e
Institutional Custody
+$2.0M
80%
0x6ef3...0476
Institutional Custody
+$1.7M
83%
0x8b3d...7e7b
Institutional Custody
+$3.5M
78%