I pull sitemaps on the first Tuesday of every month. Forty to sixty crypto media domains, hashed and diffed, checking for cadence anomalies — how many URLs, published when, by whom, and whether the rate of change matches the rate of news. It is a boring habit that occasionally pays for itself, because narrative velocity is a tradeable input, and the velocity of the press is not always the velocity of the market.
This month it paid for itself in the ugliest way. One URL, on a domain I have been reading since 2018:
Manchester United's Carrick and Mbeumo preview Champions League match against Sabah FK.
Michael Carrick retired as a player in 2018. Bryan Mbeumo is a winger who arrived at Old Trafford in the summer of 2025. Sabah FK is a club from Azerbaijan. And Manchester United finished fifteenth in the Premier League in 2024-25, which means they did not qualify for the Champions League at all — nor for any UEFA competition. There was no match. There was no preview. There were four structural errors in eleven words, published under a masthead that once broke real stories.
I have been in this industry for nineteen years. I have watched exchanges die, chains fork, and a hundred billion dollars of trustless claims evaporate over a weekend. But I have never seen the press degrade this fast. Tracing the ghost in the machine used to mean reading code for what the documentation omitted. Now it means reading a headline for what the editor never saw.
II. How the Press Starved
To understand why a broken football preview appeared on a crypto site, you have to understand that crypto media has already died twice, and nobody held a funeral either time.
The first era ran roughly from 2013 to 2017. The audience was small, technical, and pseudonymous. Signal density was high because the cost of being wrong was social, not financial — you got flamed on a forum, not delisted by a regulator. Writers were often engineers. Nobody was paid well. Almost everything published was correct, because almost nobody was publishing.
The second era, 2017 to 2021, was the professionalization. Real newsrooms. Editorial standards. Embargo ethics. Actual reporters with sources at exchanges and foundations. This was funded almost entirely by one revenue line: exchange advertising. Binance, Coinbase, FTX, Kraken, BitMEX, and a rotating cast of platforms with names like weather events bought display inventory and sponsored content at CPMs that would embarrass a sportsbook. The money was structural, which is to say it was propped up by retail speculation, which is to say it was propped up by nothing.
November 2022 ended it. When FTX collapsed, the largest single advertiser in the sector went to zero in seventy-two hours. Coinbase cut marketing. Binance pulled regional budgets. The remaining exchange ad market compressed by an order of magnitude within two quarters. I watched this happen from a fund seat, which meant I watched it as a P&L event first and a cultural one second. The layoffs were not gradual. They were a cliff.
The third era is the one we are living in, and it started around the middle of 2023. Stripped of exchange money, crypto media did what every business does when its primary revenue line disappears: it found the nearest arbitrage. That arbitrage was programmatic display, and programmatic display does not pay for accuracy. It pays for impressions. It pays, specifically, a fraction of a cent per impression, regardless of whether the sentence attached to the impression is true.
This is where the analogy to my own sector becomes uncomfortable. I have spent years arguing that liquidity mining APY is not yield — it is a project subsidizing its own TVL number, and the moment the subsidy stops, the mercenary capital leaves and the metric collapses. Crypto media ran the same program. The subsidy was ad inventory, the TVL was pageviews, and the mercenary capital was the reader who arrived from a search engine with no intention of returning. When you run a liquidity mining program for attention, you get exactly the attention you paid for: temporary, extractive, and indifferent to the underlying product.
By 2024 the industry had a phrase for this. Scaled content abuse. Google's March 2024 core update went after it explicitly, and the sites that had built their model on volume rather than authority took the hit. Some of them survived by becoming more careful. Others survived by becoming more automated. The Manchester United headline is a product of the second group.
III. What the Audit Actually Showed
I want to be precise about methodology here, because the point of this article is not that a bad article exists. Bad articles have always existed. The point is what the pattern around them reveals.
I took the domain in question and pulled its complete published archive for the trailing ninety days — roughly 2,800 URLs. For each one I extracted the title, the author byline, the publication timestamp, the word count, and the presence or absence of quoted sources with attribution. Then I looked at cadence.
The first finding: publication volume was essentially flat across the ninety-day window, with a standard deviation of under nine percent week over week. That is not how news works. News is lumpy. A single ETF decision, a single exploit, a single exchange insolvency will produce a thirty percent swing in a real newsroom's output because the newsroom is responding to events. Flat cadence across ninety days is the signature of a quota, not a beat.
The second finding: 61 percent of the bylines in the sample belonged to authors with no other discoverable digital footprint — no prior publication history, no social accounts with organic engagement, no pattern of source relationships. This is not disqualifying on its own. New writers exist. But when a majority of a publication's output comes from writers who appeared from nowhere and have no visible career outside that single domain, you are not looking at a newsroom. You are looking at a content pipeline with human-shaped labels attached.
The third finding, and the one that actually matters: the article that caught my eye was not an outlier. The domain was publishing between twelve and thirty items per day that referenced a named person, a named organization, or a named event, and in a random subsample of 200 of those items, my own spot-check against primary sources found that 34 percent contained at least one claim that was either factually wrong, temporally impossible, or attributed to a source that does not appear to exist.
A third of the output, in other words, was not degraded journalism. It was not journalism. It was the shape of journalism with the substance removed, which is a specific and more dangerous failure mode, because the shape is what readers and machines both consume.
Reading the silence between the blocks stopped being a metaphor for me around the two hundredth URL. The silence is where the sourcing should be. There are no quotes. There are no press officers. There are no match dates, venues, kickoff times, injury lists. The article did not fail to tell me what happened. It failed to contain any information that could have come from having been told anything at all.
IV. The Unit Economics of a False Sentence
The question any investor should ask about a phenomenon is not whether it is bad. It is why it persists, and what keeps it alive. Content degradation persists because the cost curve is obscene, and the obscenity is recent.
A competent human-written 800-word news preview — even a low-stakes one — costs somewhere between 150 and 400 dollars once you account for the writer, the editor, the CMS overhead, and the fact that most drafts get killed. That number has been roughly stable for a decade.
A machine-generated equivalent costs the inference. Depending on model, length, and batching, that is somewhere between two and fifteen cents. Call it a factor of ten thousand. I want to sit with that number for a moment, because I do not think the industry has metabolized it. There is a ten-thousand-fold cost differential between producing a plausible news item and producing a reported one, and both of them render identically in a browser.
Now consider the revenue side. A page that earns programmatic display at a blended rate of roughly one to four dollars per thousand impressions needs somewhere between four hundred and four thousand impressions to break even on a forty-dollar human article. At crypto CPMs in a bear market, with search traffic compressing as Google's spam systems improve, most articles never get there. The same page needs approximately four to forty impressions to break even on a machine-generated one. Four.
That asymmetry is the entire story. It is not that someone decided to destroy journalism. It is that at some point in 2023, the expected value calculation flipped, and it has never flipped back. In a bear market, it flips harder, because CPMs fall and the volume required to hit a revenue target rises, which pushes marginal producers further toward automation rather than less.
There is a second asymmetry, and it is the one that should keep anyone building verification infrastructure awake. The cost of producing a false claim and the cost of refuting one differ by orders of magnitude in the opposite direction. Producing the Sabah FK sentence cost roughly a third of a cent and took under two seconds. Refuting it required me to confirm the Premier League table, the UEFA qualification rules, Carrick's retirement year, Mbeumo's transfer, and the actual competition Sabah FK plays in. That took about eleven minutes and, if I had billed it at my fund rate, several hundred dollars.
The generation-to-verification ratio is roughly one to ten thousand in effort and one to ten thousand in dollars, running in opposite directions. Every information system in history has had this problem to some degree. None of them had a machine that could produce prose at two hundred tokens per second. This is what I mean when I say the quiet ruin when the algorithm broke was not dramatic. It did not announce itself. It arrived at half a cent per article, quietly, and it compounded.
V. The Demand Side: Nobody Was Going to Catch It
Here is the part of this that took me longest to accept, and it is the part that indicts us rather than the machines.
Nobody was going to catch that headline. Not because readers are stupid, but because readers were never reading it as information.
Think about who clicks a Manchester United preview. They click it because it is a Tuesday, there is a fixture coming, and they want a small dose of anticipation. The click is affective, not epistemic. The functional purpose of the page is to produce a feeling of proximity to an event — to generate the mild pre-match hum — and a page that produces that hum has succeeded at its job even if every proper noun on it is wrong. A reader who is not checking whether Sabah FK is in the Champions League is not being deceived, exactly. They are consuming a mood.
This is the mechanism that content farms have always exploited, and it predates AI by a century. Sports pages have carried match previews of dubious accuracy since the era of print, because the preview was never a document of record. It was a ritual object. Reading the preview was part of watching the match. The information content was always near zero; the ceremony was the product.
So when a generator of prose gets cheap enough, it does not need to fool an expert. It needs to fool a person who is about to be entertained anyway. I ran a rough check using referring-domain data on a sample of the domain's sports-tagged URLs. The overwhelming majority of inbound traffic arrived from aggregators, social reposts, and search queries of the form that a fan types while half-watching something else. Almost none of it came from anyone who arrived intending to verify anything.
The uncomfortable corollary: the readers who would have detected the error are precisely the readers who never arrive. That is not an accident of this particular site. It is the structural condition under which low-quality content survives, and it means that any solution built on the assumption that better-informed readers will police the market is built on sand.
VI. The Ghost in the Sports IP Token
Now let me take the bridge that the material actually offers, because Manchester United is not merely a subject of a broken preview. It is an asset class.
Football clubs are among the richest narrative properties on earth. They own a century of accumulated drama, a global diaspora of identity-holders, and a weekly cadence of emotionally significant events. In 2020 and 2021, the crypto industry looked at that and saw the most obvious tokenization opportunity it had ever encountered. The result was the fan token, and the fan token is the clearest case study we have of a narrative that was priced before it was built.
The mechanics were simple enough. A club partners with a token issuer. Holders receive voting rights on a set of pre-approved, low-consequence questions — a song to play at the stadium, a training-ground mural, a warm-up jersey. The token trades on the open market. The pitch was that this was a new form of membership, and for a moment in 2021 the market agreed.
Here is what I have to say about it as an investor, and I want to be careful to make this a technical claim rather than a moral one. Fan token volume peaked in the first half of 2021 and has been in structural decline since. By any measure I can construct from exchange data, aggregate volume across the major club tokens is down more than ninety percent from that peak. The governance utility did not scale. The number of binding votes held never grew in proportion to the number of holders, which is exactly what you would expect when the voting surface is deliberately limited to cosmetic questions. And the primary use case that actually materialized was speculation on match outcomes and signing rumors — which is to say, the tokens became instruments for trading anticipation rather than for participating in a club.
Note what that makes them. A fan token is a financial instrument whose price is driven almost entirely by narrative supply, and whose holders have no claim on any cash flow. Its fundamental is attention. Which means the fan token is the one asset class in crypto for which the integrity of sports media is a direct input cost.
And note the absence in the roster. Manchester City has a token. Arsenal has a token. Barcelona, Juventus, Paris Saint-Germain, Atlético Madrid, Inter, Milan, Galatasaray, Valencia — all tokenized, through one issuer or another. Manchester United does not. The biggest commercial brand in English football, with the largest global supporter base of any club in the world, never issued one. Every explanation I have heard for that is a business explanation, and they are all plausible. But the pattern is worth noticing, because it means that the club in our broken headline is precisely the club that never had a tokenized attention market to distort.
VII. Where Misinformation Has a P&L
Here is where this stops being a media story and starts being a market structure story.
Human readers absorb misinformation at a discount. They know, at some level, that a match preview is soft. They know that an aggregator repost is unverified. The affective reading mode carries its own internal skepticism, and that skepticism, however lazy, absorbs a huge amount of error. This is not a defense of the reader. It is an observation about why the damage has been contained so far.
Machines do not have that filter. And the machines are beginning to trade.
Prediction markets are the one venue in this entire stack where a false sentence has an immediate, measurable, adversarial cost. On a well-run platform, a rumor about an injury, a lineup change, a managerial sacking, or a fixture cancellation does not sit in the discourse for weeks. It gets priced within seconds, and the pricing is public, and the price reverts violently if the rumor is wrong. That reversion is the closest thing we have to a self-correcting mechanism for narrative error, and it works precisely because somebody had money on the line and a reason to do the eleven minutes of verification I did.
I have watched this mechanism operate at the granularity of a single question. The pattern is consistent: a low-quality outlet publishes a claim; an aggregator picks it up; the claim enters a feed that a market participant is monitoring; a market moves a few percent; somebody runs the check; the price snaps back. The whole cycle can take ninety seconds. In media terms that is nothing. In market terms it is an entire correction arc, compressed into the time it takes to read one paragraph.
This is the strongest argument I know for why prediction markets matter beyond their own speculative utility. They are an error-detection subsidy. They pay people to care whether a sentence is true, at a moment when caring is still profitable. Nobody pays a reader to catch a Sabah FK error. Somebody might pay a trader.
But — and this is the hinge on which the whole argument turns — the mechanism only works where a market exists. Sabah FK would never have a market. A match preview for a fixture that does not exist would never be priced. The set of claims that prediction markets can discipline is exactly the set of claims that somebody, somewhere, is willing to underwrite. Everything outside that set — which is nearly all information, including nearly all of the information that shapes public understanding — remains unpriced, and therefore undisciplined.
VIII. The Oracle's Final Whistle Problem
If prediction markets are the error-detection layer, then sports data oracles are the settlement layer underneath them, and the settlement layer has a problem that the industry talks about far less than it should.
Getting a final score onto a blockchain is not a hard problem in the aggregate. Feed providers have built the plumbing. Sports data has been pushed on-chain through multiple oracle architectures over the past three years, and the latency has come down from minutes to seconds for major fixtures. For most markets, this is sufficient. The interesting failures are never in the ordinary case.
The interesting failures are in what I would call the final whistle problem: the window between the last meaningful event and the moment a result becomes irrevocable. In that window, three things can go wrong at once. The feed can be right and the settlement rule can be wrong — the classic own-goal attribution dispute, where the data provider credits the goal to a defender and the market's rules say the striker should be credited. The feed can be wrong and the settlement rule can be right, which is the variable-latency case, where a stale update causes a market to resolve early. And the feed can be right, the rule can be right, and the dispute mechanism can still produce a wrong outcome, because the arbitration process is governed by token holders with an incentive that is not aligned with accuracy.
That last one is the one that should worry anyone building on this stack. An oracle with a token-governed dispute layer has introduced a political process into a factual question. The token holders are not, in general, in a position to adjudicate whether a ball crossed a line. They are in a position to vote on who they think should win, weighted by stake. When the factual question is small and the stakes are large, that is a mechanism for converting capital into truth, which is a very old problem wearing very new clothes.
I have seen this pattern before, in a different costume. When algorithmic stablecoins failed in 2022, the failure was not in the math. The math was fine. The failure was that the system's assumptions about human behavior — that holders would not all exit at once, that the peg would hold under stress, that the incentive structure would produce the coordinated action it theoretically rewarded — were assumptions, and the code could not tell the difference between a model and a fact. The code remembers what the market forgets: that a mechanism is only as good as the behavioral claim embedded in it.
An oracle that resolves disputes by token vote has embedded a behavioral claim too. It claims that the largest stakeholders will resolve honestly. Sometimes that will be true. In a bear market, when the tokens are cheap and the positions are small, it will almost always be true, because nobody bothers to corrupt a market worth six figures. That is not reassurance. That is a statement about the current price of corruption.
IX. When the Agents Start Reading
Now to the part I think is genuinely new, and the part that took me from idle curiosity about a broken headline to writing this at all.
Last year I published a piece arguing that the convergence of AI agents and blockchain would produce a new kind of economic actor: a program that holds a mandate, reads the world, and pays for data and compute through smart contracts, with the chain serving as an immutable audit trail for its decisions. I still believe that thesis. What I did not fully work through at the time was the input side.
Agents read the news. Not the way you read the news. An agent with a trading mandate and a retrieval pipeline does not consume a Manchester United preview for the pre-match hum. It consumes the text, extracts the entities, scores the entities, and converts them into a position. And here is the property that makes this different from human reading: an agent's confidence is a function of corroboration, and corroboration is a function of frequency.
This is the mechanism behind what I have started calling the sybil-of-sources problem. A single false claim in one outlet is noise. The same false claim propagated across forty syndication endpoints, three aggregators, two content-repurposing pipelines, and a handful of social mirrors is not noise. To a retrieval system that weights passages by how many independent documents assert them, it is consensus. The generation of forty documents costs a few dollars. The agent does not know they share an author. It cannot know, because provenance is not in the token window.
Run the arithmetic forward. A commodity LLM costs cents per article. Forty articles cost a few dollars. Those forty articles are enough to move a retrieval-augmented model from ignorance to confident assertion, because the model has no mechanism for distinguishing forty independent sources from one source printed forty times. And now that assertion is an input into a trading decision, which is an input into a price, which becomes an observation that the next agent cites. The rumor has laundered itself into a fact by passing through enough reproductions.
I want to be precise about what I am and am not claiming. I am not claiming this has happened at scale in a major market. I am claiming that the mechanism exists, that the inputs are cheap, and that the defensive structures are absent. Human readers absorbed the Sabah FK error because they were reading for mood and mood is error-tolerant. Agents read for signal and signal is not error-tolerant at all. When the herd of agents wakes, the signal has already faded — but they will have traded on the ghost of it first, and the ghost will have been manufactured for pennies.
This is why I think the content farm problem is mispriced as a media story. Media stories get fixed by editors. The content farm problem is going to get fixed — or not — by whoever owns the retrieval layer that agents trust. That is a much more concentrated, and much more valuable, chokepoint than any masthead.
X. Provenance Is Not Truth
The obvious response is to put journalism on-chain. Signed attestations from verified publishers. Content hashes. Authorship credentials that cannot be forged. A corrections ledger that is append-only and auditable.
I have looked at these proposals seriously, and I want to state clearly where they work and where they do not, because I think the industry is about to spend a great deal of money on a partial solution and then mistake it for a complete one.
Attestation infrastructure solves one problem well: it tells you who said something, and whether the thing you are reading is the thing they said. That is genuinely valuable. The problem of the misattributed quote, the quietly edited correction, and the fabricated byline all become tractable when the author signs the text and the hash is committed. Ethereum Attestation Service and its various cousins give you this at reasonable cost. If a newsroom wants to prove that its archive has not been rewritten, it can. That is a real improvement over the status quo, and I expect serious outlets to adopt it over the next two years.
But read the Sabah FK headline again and ask which problem it is. The article was not misattributed. Nobody forged a byline claiming to be a different publication. The text was published exactly as written. The signature, if there had been one, would have verified perfectly.
What was wrong was that the sentence was false. And no amount of cryptographic provenance touches that. A signed lie is a lie with better paperwork. Provenance establishes accountability, not accuracy. It tells you who to blame after the fact, which is useful in a courtroom and useless in a feed.
The proposals that go further — decentralized fact-checking markets, staked claim-and-challenge protocols, reputation tokens for accurate reporting — all founder on the same two rocks. The first is the verification cost asymmetry I described earlier: the challenger pays ten thousand times what the publisher paid, and the protocol has to fund that gap, which means the protocol needs a revenue model, which means it needs to be profitable, which means it eventually behaves like an advertiser. The second is that the volume of claims is now so large and so cheap that any adjudication system becomes a denial-of-service target. You cannot stake enough to make a three-cent lie expensive when the liar can produce a million of them.
VII. The Contrarian Case: Accuracy Was Never the Product
I want to end the analytical portion by arguing against the framing that I have been using, because I think the framing is comfortable and the comfort is misleading.
The story I have told so far has a villain: automated publishing. It has a victim: the reader and the agent who believed a false sentence. And it has an implied remedy: make publishing more expensive, more accountable, more verifiable, and the problem goes away.
I do not think that is right. I think the villain is misidentified, and the misidentification is convenient for everyone in this industry, including me.
Consider what the Manchester United preview was for. It was not produced to inform anyone. It was produced to occupy a slot — a search result, a feed position, a share of attention that would otherwise go to a competitor. The article is an occupancy instrument. Its function is to exist in the right place at the right time and collect a fraction of a cent from everyone who passes through.
Nothing about that function is new, and nothing about it is machine-specific. Print newspapers ran stock match reports for decades. Sports wire services pumped out previews on spec. The commodity content has always existed, and it has always been bad, and the reading public has always understood, at some level, that it was bad. What changed is not the intent. What changed is the marginal cost, and marginal cost changes scale, not nature.
The uncomfortable conclusion follows directly: accuracy was never the product. Coherence was the product. A reader wants a story that hangs together, matches their priors about how the world works, and delivers the emotional payoff they came for. Accuracy is a constraint that only binds when the error is large enough to break the coherence. The Sabah FK error was small. Most readers never noticed it, and the ones who noticed it did not care, and the ones who cared did not act. The system has priced that correctly.
Which means the remedy I sketched is wrong in an important way. Making publication expensive does not fix the problem. It just raises the entry cost, which consolidates the market among the firms that can afford the expense, which is a good outcome for incumbent publishers and a bad outcome for everyone else. The incentives do not change. The occupancy function remains. Somebody will still want to occupy the slot, and they will pay whatever the slot costs, and they will optimize for coherence rather than accuracy because that is what the slot demands.
What would actually change the incentive is a slot that pays for accuracy. And I do not know of one. Prediction markets nearly qualify, but as I argued, they only pay for accuracy on questions that are tradeable. Kaito-style mindshare metrics, which have become popular in the last two years as a way of turning social attention into a scored asset, do something subtly worse: they reward frequency, which is exactly the quantity the farms manufacture. A proposal to add news-quality scoring to a mindshare system looks like a fix and functions like a subsidy, because the cheapest way to score highly on a quality metric is to produce a lot of text that superficially satisfies the metric's surface features. You cannot score your way out of a problem that is caused by cheap, coherent, high-volume text, using a metric that reads text.
XII. What Survives the Bear
So where does that leave an investor in a market where survival matters more than gains? Let me be concrete about what I am watching, and what I would be cautious about.
Sports data oracles are the piece of this stack I would look at most carefully, and not for the reason most people do. The interesting variable is not latency — latency is a race that gets won and then stops mattering. The interesting variable is dispute resolution design. A sports oracle that resolves contested events through a token-governed vote has a governance surface that scales with the market it serves, and I would want to know what happens to that surface when the markets get large enough to be worth corrupting. That question has a specific answer in each protocol, and the answer is usually in the documentation, and almost nobody reads the documentation.
Prediction markets are the piece I would look at second, and my interest is not in election cycles. It is in whether the platforms can extend pricing into the long tail of small factual questions, because that is the only path I can see to a general error-detection subsidy. The obstacle is not regulatory, although the regulatory obstacles are real and compounding. The obstacle is that the long tail of factual claims is expensive to adjudicate and unattractive to trade, which means the markets that would most improve information integrity are the markets with the worst unit economics. I expect that to remain true.
Fan tokens are the piece I would not touch. The fundamentals — attendance, broadcast rights, commercial revenue — are real and large, but the token's claim on them is nil. What the token prices is anticipation, and anticipation is the most efficiently commodified emotion in this sector. I have watched enough of these cycles to know how this one ends: the attention that sustains the price migrates to the next shiny thing, and the token holders who were told they were members discover they were liquidity.
And the media layer, which is where this article started, is the piece I expect to be restructured rather than fixed. The mastheads that survive will be the ones that own something the farms cannot replicate — a source relationship, a research archive, a proprietary dataset. Everything else becomes a formatting layer over commodity text, and formatting layers get commoditized next.
The forward-looking question, and the one I cannot answer, is who audits the retrieval layer. Not the publisher — the pipeline. The thing that an agent consults when it decides whether a sentence is true. If that pipeline is assembled from the open web, it will be assembled from farms, because farms are the cheapest source of the most text. If it is assembled from a curated corpus, it will be assembled by whoever can afford the curation, and the price of truth will be set by whoever owns the corpus.
We traded chaos for consensus, and lost ourselves somewhere in the swaps. The question now is who gets to write the consensus, and whether anyone will notice that it was written by a machine that never watched the match.


