The Verification Layer
The gap between AI hype and production reality
The viral pitch says twelve prompts replace a quarter-million-dollar analyst. It is correct about the frameworks and wrong about everything that pays the analyst. A note on what AI made cheap, and what it quietly made more dangerous.
A post is moving through my feed at the moment. You have probably seen its relatives. The headline is a variant of the same trade: Wall Street’s $250,000 analyst stack, reduced to twelve copy-paste prompts and one $20 subscription. Bloomberg-grade research at one-thousandth of the cost. A thirty-hour analyst week compressed into twenty-five minutes.
It arrives with an elegant infographic. A four-level system. A numbered workflow. A gravestone reading Bloomberg, 1982–2024, killed by Claude. It is well made. It is also, in the precise sense, a mispriced instrument, and it misprices itself in exactly the way I spend my working life documenting in other people’s balance sheets.
Let me take the argument seriously before I take it apart, because part of it is correct.
What the pitch gets right
The claim runs like this. The frameworks institutional analysts use were never proprietary. Discounted cash flow, comparable-company screening, scenario analysis, the standard risk methodologies – all of it sits in CFA study guides and MBA curricula, available to anyone who wants it. The intellectual property was never the secret. The moat was execution. And AI has now collapsed the cost of execution to almost nothing.
The first half of that is true, and it is worth conceding plainly. There is nothing secret in the modelling. A competent graduate can build a three-scenario DCF from a template. The methods are public and have been for decades. The pitch is right that nobody was ever paying for the frameworks.
It is the second half, the claim about execution, where the whole thing comes apart.
What execution means on the buy side
The pitch treats execution as the production of the artifact: the finished model, the formatted report, the populated screen. On that definition, the cost has indeed fallen to near zero. A system that returns a fluent five-section equity report in ninety seconds has clearly collapsed the cost of producing five-section equity reports.
But producing the artifact was never the expensive part. An Excel template produced the artifact in 1998. The junior analyst’s £150,000 seat is not funded to build the DCF. It is funded to know which input is wrong. To notice that the revenue assumption quietly compounds at a rate the company has never once achieved. To catch the off-balance-sheet item the model omits because the template had no row for it. To recognise when a management team is laundering a weak quarter through a change in segment reporting.
That is execution on the buy side. Not generation, adjudication. Knowing which of the plausible numbers in front of you is the one that breaks the thesis.
AI collapsed the cost of generation. It did not touch adjudication. If anything it raised it, because the output now arrives fluent, confident, and formatted to look precisely like the work of someone who checked. The expensive judgement did not disappear. It became harder to see that it was missing.
The category error
There is a deeper problem, and it is technical rather than rhetorical.
A Bloomberg terminal is deterministic. You query a figure and you receive the same figure every time, sourced and reproducible. A sell-side note is a fixed object; it says what it says. These are instruments you can build on, because their outputs do not move when you are not looking.
A large language model is not that kind of object. It is a sampling process over a probability distribution. The same prompt, run against the same company on two different days, will not return the same model. Phrasing shifts the output. Context-window position shifts the output. Temperature shifts the output. This is not a defect awaiting a patch; it is the nature of the substrate.
Which is why one line in the pitch should stop any practitioner cold. The promise of “consistent output, no rewriting” is not an overstatement of a real feature. It is a description of a property the technology does not possess and cannot possess. You cannot pour a foundation on a process whose outputs you are unable to reproduce. The claim is not ambitious. It is a category error, infrastructure language applied to a probabilistic text generator.
The rebuttal is printed on the same page
Here is the part that should settle it.
The infographic carries its own fine print. A golden rule, set off in the corner: cite every number, verify every claim, challenge every assumption. Elsewhere, a section on the six mistakes that produce weak outputs and the fixes that correct them.
Read that against the headline. If the system produced consistent, Bloomberg-grade output from twelve copy-paste prompts, it would not ship with a troubleshooting guide for weak outputs. It would not need a standing instruction to verify every number it gives you. That instruction is not a flourish. It is the entire job, smuggled into the footnotes and set in a smaller font than the promise it contradicts.
Work it through. If you must cite, verify and challenge every figure across twelve sequential prompts, you are not running a twenty-five-minute workflow. You are running the same forensic audit the analyst ran – by hand, without the analyst’s training, against output engineered to look as though it has already been checked. The headline sold you an autopilot. The fine print hands you a shovel and tells you to dig.
The shape of the thing
I find this structure familiar, because it is the structure of nearly every mispriced security I am asked to examine.
There is always a clean headline number. A treasury company trading at a flattering multiple to the assets it holds. A premium narrative with a single figure doing all the persuasive work. And there is always a second document, the credit facility terms, the concert-party disclosure, the note at the back of the filing, where the risk the headline omits has been parked: technically disclosed, practically unread.
The forensic reader’s discipline is simple and unglamorous. Ignore the headline; audit the footnotes. The number on the poster is the thesis the model was built to prove. The liability is wherever the smaller type is.
“Bloomberg-grade for $20” is the headline. “Verify every claim” is the footnote. They are the same instrument, and the instrument is mispriced.
The contradiction already in print
One further observation, offered without naming anyone, because the pattern is the point.
The same operator selling single-pass prompt-chaining to the open feed maintains, behind the paywall, a post advising paying subscribers that prompt-chaining is already obsolete, that the frontier has moved on. Two theses, mutually exclusive, priced at different tiers. The free audience receives the dream. The paying audience receives the walk-back.
This tells you what the product is. It is not a research system. The frameworks, as the author rightly says, are public; you could assemble the prompts in an afternoon. What is being sold is certainty. And certainty is the one output the underlying technology is structurally incapable of supplying. The margin lies in manufacturing the confidence the tool cannot.
What rigour looks like
For contrast, the method on this desk.
A recent report specification ran to sixteen versions. An agent build, nine. Not because the tool is poor, but because every “this is complete” from a model is the start of the verification step, not the end of it. Find the errors. Action the feedback. Run it again.
And no single instance is trusted to check itself. A live specification is being red-teamed across four model families in parallel – three Claude instances, DeepSeek, Gemini, Qwen. Different architectures fail in different places, and the points where they disagree are precisely where the defects sit. Cross-model disagreement is signal. A single model’s fluent confidence is not. Running one model, single-pass, and shipping its first answer is not efficiency. It is the absence of a control.
The spec is still defective, after all of that. That is not the embarrassing admission the pitch would have you treat it as. It is what verification looks like when the output has to hold in front of an institutional reader.
The cost that did not collapse
Strip it back. AI did not reduce the cost of analysis. It reduced the cost of producing something that looks like analysis, and those are not the same good. One is an asset. The other, unaudited, is a liability formatted to pass for one.
The expensive part of the work was never the model. It was knowing when the model is wrong. That skill did not get cheaper. It became more valuable, and considerably more rare, at the precise moment the tools learned to produce confident, fluent, plausible error at scale.
Somewhere right now, someone with no engineering background and complete confidence is shipping one of these single-prompt systems into a context where the numbers matter.
God knows what it does.

