Null Input, Confident Output: The Research Pipeline Leak Nobody Prices
The Document
A 2,400-word analyst report crossed my desk last week. Nine sections. Four scorecards. A risk matrix with six categories and five columns. An information-value table. A composite verdict. A disclaimer.
Every field read N/A.
Not "insufficient data." Not "pending verification." The literal string N/A, repeated across more than a hundred cells, wrapped inside a schema that presupposed the existence of things that did not exist. The document assigned zero stars across technical value, investment value, timeliness, and reference value. Then it rendered a disclaimer and signed off.
The analyst did nothing wrong.
That is the part worth your attention. The pipeline executed exactly as designed. Input arrived empty. The system still produced a deliverable, because the system was never built to produce the alternative, which is silence.
I have spent eleven years watching crypto infrastructure confuse output for signal. This is the cleanest instance I have seen, and the most dangerous, because it scales.
Where The Leak Is
Roughly a third of the "research" moving through crypto Telegram and X right now is generated by a two-stage pipeline. An extraction model reads a source. A synthesis model writes nine sections about the extracted facts. When extraction works, the output is mediocre but harmless. When extraction fails, you get a structurally flawless document with zero information content.
The failure is not the model's fault. The failure is the schema's fault.
The synthesis prompt carries a fixed taxonomy. Technical analysis. Tokenomics. Market structure. Ecosystem positioning. Regulatory exposure. Team and governance. Risk matrix. Narrative. Supply-chain transmission. Nine buckets. The model is instructed to fill all nine. It is not instructed to refuse.
So when extraction returns an empty fact list, the model does not stop. It cannot stop. It has been handed a container and told the container must be full. The cheapest way to fill a container is with declarations of emptiness. N/A is a token. It costs nothing to emit. It satisfies the schema.
I have seen this exact pattern in a different domain. In early 2025 I built an API wrapper to trade against a handful of emerging AI-agent bots on decentralized venues. The bots were marketed as autonomous market participants. Functionally, they were schema executors. They had been told to react to volume spikes. They reacted to volume spikes — regardless of whether the spike carried information.
A wash trade from a single wallet bouncing between two addresses it controls looks identical to organic flow if your only input is the volume number. The bot did not care. It had a rule. The rule said buy. So it bought.
By the third week I was running 150-plus counter-trades per day against that behavior, with a 58% hit rate and roughly $42,000 in monthly gross. Not because I was smarter than the bots. Because I read the input they were ignoring.
Code is law, but math is the judge. The bots compiled. The math did not.
Three Stages, One Leak
"Bad AI research" is not a useful diagnosis. There are three distinct stages in the pipeline, and only one leaks.
Stage one: ingestion. Fetches the source. Cheap. Reliable. Rarely fails.
Stage two: extraction. Converts unstructured text into structured facts. This is the expensive stage, and it fails silently. An extraction model that finds nothing returns nothing. No error code fires. No exception is thrown. An empty list is a valid return value.
Stage three: synthesis. Renders facts into the required schema. This stage is nearly impossible to make fail, because a sufficiently capable language model can generate fluent prose about the absence of facts. That is one of the things it is best at.
The architecture has a single point of failure in the middle and no circuit breaker on either side. Ingestion does not validate extraction. Synthesis does not validate extraction. The output reads as authoritative and carries zero bits of transferable information.
Now map that onto a trading desk.
A desk has the same three stages. Data ingestion. Signal extraction. Position synthesis. The difference is that on a desk, a failed extraction produces a P&L number, and the P&L number does not care how good the document looked. This is why I trust a fill log more than a research note. The fill log cannot lie to me. It has no schema to satisfy.
What Real Extraction Looks Like
Let me be concrete about the work that the null report was pretending to do.
In late 2023 I spent roughly 200 hours reverse-engineering Lido's stETH rebalancing mechanism on-chain. Not reading the docs. Reading the contract. Tracing oracle updates across blocks, reconstructing the accounting state after each report, and replaying the sequence under simulated congestion.
What I found was a reentrancy surface in the oracle feed that opened during high network load. I reported it through the official bug bounty channel. It paid $5,000.
That entire effort produced exactly one finding. One. Not nine sections. Not a four-row scorecard. A single deliverable that fit in a paragraph.

This is the arithmetic the research pipeline never runs. Real extraction is expensive, slow, and produces a low count of high-value facts. Schema-driven synthesis is cheap, fast, and produces a high count of low-value strings. When you reward the second and bill for the first, you get null reports.
The Lido exercise also changed how I read yield. Every APR is a price. The price is being set by someone who knows the contract better than the person buying the yield. Code-level skepticism is not a personality trait. It is a discount rate.
Sideways Markets Punish Fake Signal Hardest
We are in chop. That matters more than usual here.
In a trending market, a bad research note is diluted. Direction does most of the work. If everything is going up, a report that says nothing still gets attributed to a correct call.
In a range, there is no direction to hide behind. Positioning is the entire game, and positioning requires an actual edge. Chop is where fake signal gets priced. Not immediately. But reliably.
Here is what I watch when I need to distinguish real extraction from synthesized emptiness. Four measurements, all observable, none of them requiring an opinion.
Bid-ask spread. Not the headline number on the aggregator. The realized spread on the pairs you actually trade. When liquidity thins, spread widens before price moves. A widening spread with flat price is a positioning signal, and it is not in anyone's nine-section template.
Perp funding. Neutral funding with rising open interest means longs and shorts are both paying to hold. That is a coil. Not a direction. A coil.
Gas priority fees and mempool composition. If priority fees are climbing while on-chain volume is flat, someone is paying for inclusion. That someone is usually not retail.
Depth-of-book asymmetry. Equal volume on both sides of the mid is rare. When it appears, the book is being managed, not cleared.
None of these require a narrative. All of them are extractable. None of them will appear in a schema-driven report, because none of them are in the schema.
The MEV Layer Nobody Prices
There is a second leak, and it runs in the same direction.
In 2020 I was running Python scripts against the Ethereum mempool, watching for large Uniswap V2 swaps. I executed 47 arbitrage swaps across SUSHI and 0x over three weeks for roughly $12,400 gross. The lesson was not that arbitrage exists. The lesson was that the value was captured before the trade printed. By the time the swap appeared on a chart, it was already history.
DEX aggregators now market "best route" execution to retail on the promise of saved fees. I have looked at the routing tables. The fee saved is real and small. The value extracted by searchers sandwiching the resulting flow is real and large. The aggregator optimizes the price it shows you. It does not optimize the price you get.
That gap is a leak, and it is the same class of leak as the null report. Output is being manufactured to satisfy a display schema. The display schema is "best route." The reality is a spread.
Code is law, but math is the judge. The math says the aggregator's savings and the searcher's extraction are drawn from the same pocket. One of them is disclosed. One is not.
The Institutional Layer, Briefly
One more data point, because it is the cleanest example I have.
After the BTC ETF approval in January 2024, I ran a cash-and-carry book. ETF share price against the underlying futures basis, sized to the observed convergence. Six months, $250,000 notional, 3.2% annualized, roughly $8,000 net. Risk-free in the accounting sense, not in the metaphysical sense.
The relevant observation is not the return. It is that institutional entry did not remove the inefficiency. It changed the counterparty. The basis existed because two different populations wanted two incompatible things — one wanted exposure, one wanted financing — and the market charged a toll for connecting them.
I stopped chasing narrative pumps around that time and started chasing structural inefficiencies. Narratives are schema-driven. Structures are measured.
The Contrarian Read
Everyone is looking at the null report and concluding that AI research is unreliable. That is the wrong conclusion, and it is the one the pipeline wants you to reach.
Here is the correct one. Retail reads confidence as competence. Smart money reads absence of signal as signal.
A document that confidently reports zero findings, in a schema built for nine, is not a broken document. It is a truthful one wearing a costume. The costume is the problem.
The deeper contrarian point: the same pipeline that generates empty research now feeds a growing share of automated trading activity. The bots I traded against in 2025 were downstream consumers of exactly this kind of input. When extraction fails upstream, the emptiness does not stay upstream. It propagates as noise-driven order flow, which manufactures volume, which other bots read as signal, which manufactures more volume.
That is not a market. That is an amplifier with a fee schedule.
And here is what the schema cannot produce, ever: the phrase "no trade today." There is no column for it. A pipeline paid per word cannot output silence. So it outputs N/A instead, and the reader mistakes the volume of text for the volume of information.
Code is law, but math is the judge. The math here is simple. A document with zero bits of information has zero expected value, regardless of its length. The market has not priced that yet, because the market is still reading word counts.
Takeaway
In a range, the highest-EV position is frequently the one you do not take. That is not a philosophy. It is a statement about variance.
What I am watching for over the next two quarters is whether anyone prices the leak. Concretely: research products that bill for negative results. Extraction pipelines with hard validation gates between stage two and stage three. Bots that are rewarded for not trading.
That last one is the tell. When you see a market participant that gets paid for inaction, you are looking at real extraction. Everything else is a schema looking for something to fill.
Around here, the spread will tell you first.