A single capability, contract review, can now be built in more distinct ways than a diligence team could test in a quarter. The model underneath can be any of a dozen frontier systems. Each can be aimed through prompting, retrieval, fine-tuning, or tool use. Each sits inside a product that wraps it in workflow.
Independent benchmarks exist, but they measure models in isolation, not the product a buyer actually pays for. The distance between the two is where advantage is won and lost.
Every reported measure of an AI business blends two different things. One is the market it sells into: demand, competition, switching costs, regulation. The other is the product itself: what it can actually do, and how hard that is to reproduce.
The market side is well served. Customer conversations, channel checks, and comparable analysis read it reliably. The product side is the harder one, and it has become harder still. What a model can and cannot do is now changing on a timescale of months.
The product fundamentals now shift month to month as models improve. That is the half this note sets out to measure.
Reading the product side becomes tractable once the product is broken into parts. Whatever the application, the same four components are present, and each can be measured on its own.
A compound, a product's overall performance, tells you nothing about its individual parts. To read one component, hold everything else fixed, move only that one component, and observe the change in outcome.
This is the standard method of experimental science. It is the one thing that assessment based only on the outward result can never do, because from outside you observe only the combined effect. The moment a component can be isolated and varied independently, its contribution stops being a matter of opinion and becomes a measured fact.
Each component has its own way of being tested, but the method is the same: isolate it, vary it, and read the result against independent ground truth.
The legal-technology study in the following sections is one applied example of this general method.
A base model is now the default unit of capability. It is the basic competence a product gets out of the box, before any wiring, data, or tuning. The quality of the baseline unit rose sharply, and the price for equivalent quality fell by roughly ten times a year. A capability that cost about sixty dollars per million tokens in 2021 costs a few cents today.
Figma's strong unit was a design-surface capability. Atlassian's was a workflow capability. Adobe's was a file-format and editing capability. Baseline generative models began producing and editing those artefacts directly. Adobe's forward multiple fell from roughly thirty times to twelve.
These were well-engineered companies with strong products and real customers. What they shared was proximity to the baseline unit, and that is what drove the repricing. A product whose advantage sits in data or connections is a separate case.
The collapse in the baseline price does not make units worthless. It moves the value to units that a competitor cannot simply rent. A unit regains durable value in three main ways. Each turns the unit from a shared commodity into something specific to the product, and each is measurable.
The question for any product built on a model is simple: is its unit still the generic one, or has it become one of these? The answer sets how exposed the product is to the next model release.
A unit trained on data no one else has (proprietary weights, not just architecture) stops being generic.
A model tuned for a narrow domain (tabular data, a legal corpus, a modality) can beat a larger general one on that domain, at lower cost.
On-device, in a secure environment, or under latency and cost limits a frontier model cannot meet.
Data helps until it stops helping. Adding examples improves a product quickly at first, then less, then barely at all. The point where the curve flattens is where the data advantage is effectively spent.
Where that point sits differs by product. A product with genuine compounding infrastructure (an email-deliverability system, where more volume keeps improving performance) keeps climbing. A product whose data advantage is shallow flattens within a handful of examples. The plateau location is specific to each product and has to be measured, not assumed.
We took one contract-review product, the kind seen in a legal-technology deal, and measured how much its performance improved as we supplied more worked examples, scoring each result against answers expert lawyers had marked correct.
The first five examples moved the score by 0.286. The next forty-five moved it by 0.041. The advantage was front-loaded and then flat. Past a handful of examples, more data bought almost nothing.
This is the third and fourth component read on the same product from the same study. With the unit and the available data held fixed, we rebuilt the work one layer at a time. Retrieval added nothing. Definitions added nothing. The parts that looked like the engineering moat moved the score barely at all.
The parts that looked like the engineering moat moved the score barely at all. What appears to be defensible plumbing often is not. Each layer has to be tested on its own before it can be credited with the advantage.
Three products Caliper Lab has evaluated, anonymised here and read through the same four-component method. The same analysis places the advantage in a different layer each time, and rarely where the pitch would point.
A library of completed reads across products and sectors, each mapping where a real asset's advantage sits.
Accumulated measurement runs across models, tasks, and configurations, so each new read starts from evidence, not from zero.
Standing infrastructure that re-runs evaluations as models change, tracking how fast a baseline is closing on a given product.
In each case the question is the same: what actually carries this product's performance, how durable is it, and what would a competitor have to reproduce.
Caliper Lab is an independent AI research and evaluations firm. We read the product itself, component by component, against independent ground truth.
Built by a team of AI researchers, scientists, machine learning practitioners, and domain leaders.
An outside read, with no stake in the products we measure.
Every number traces back to an independently constructed task and a logged run.
Continuous runs track how fast the competitive baseline is closing.
We are actively expanding our library of use cases and evaluations. If you have a product to assess, or a question about how we work, we would be glad to talk.