Caliper Lab · Research note

Assessing AI,
when you cannot
see inside it.

Scroll to explore
The measurement gap

The number of ways to build an AI product has outrun the market's ability to evaluate them.

A single capability, contract review, can now be built in more distinct ways than a diligence team could test in a quarter. The model underneath can be any of a dozen frontier systems. Each can be aimed through prompting, retrieval, fine-tuning, or tool use. Each sits inside a product that wraps it in workflow.

Independent benchmarks exist, but they measure models in isolation, not the product a buyer actually pays for. The distance between the two is where advantage is won and lost.

Exhibit 1 · The stack, three tiers
Product
The UI/UX layer wrapping the model and the harness.
Dashboard
Workspace
Reports
Settings
Harness
How the model is aimed: prompting, retrieval, fine-tuning, tool use, context.
Prompting
Retrieval
Fine-tuning
Tool use
Context
Model
The base capability, now rentable.
OpenAI Anthropic Google Meta xAI + more
Where the difficulty sits

Two forces set an AI product's value, and only one of them can be read from the outside.

Every reported measure of an AI business blends two different things. One is the market it sells into: demand, competition, switching costs, regulation. The other is the product itself: what it can actually do, and how hard that is to reproduce.

The market side is well served. Customer conversations, channel checks, and comparable analysis read it reliably. The product side is the harder one, and it has become harder still. What a model can and cannot do is now changing on a timescale of months.

Exhibit 2 · What a value read is made of
MARKET demand · competition switching · regulation Readable from outside PRODUCT what it can do how hard to reproduce Changing monthly The reported metric blends both.

The product fundamentals now shift month to month as models improve. That is the half this note sets out to measure.

The frame

Any software product is a combination of four components.

Reading the product side becomes tractable once the product is broken into parts. Whatever the application, the same four components are present, and each can be measured on its own.

Units
The functional blocks that do the work (a model, a classifier, a solver). A model is now the baseline unit, available to rent.
Data and knowledge
Everything the product carries: raw data, structured records, encoded know-how (templates, labelled sets, domain rules).
Connections
How the parts are wired together and to outside systems (retrieval, integrations, APIs, workflow steps).
Logic
How the model is aimed and orchestrated (prompts, decomposition, tool routing, control logic).
Units
Data & knowledge
Connections
Logic
The isolation principle

To understand a product, isolate each component, test it, and read it on its own.

A compound, a product's overall performance, tells you nothing about its individual parts. To read one component, hold everything else fixed, move only that one component, and observe the change in outcome.

This is the standard method of experimental science. It is the one thing that assessment based only on the outward result can never do, because from outside you observe only the combined effect. The moment a component can be isolated and varied independently, its contribution stops being a matter of opinion and becomes a measured fact.

Hover a row: three components hold, one moves
Exhibit 3 · Hold the rest. Move one. Read the outcome.
Units Held
Data and knowledge Held
Connections Held
Logic Held
Outcome
0.34
Four components, one method

Four components, four dimensions, one principle.

Each component has its own way of being tested, but the method is the same: isolate it, vary it, and read the result against independent ground truth.

Units
How much the unit adds over the best available rented baseline, and how quickly that baseline is catching up.
Data and knowledge
The marginal value of each additional increment of data, typically steep at first, then bending, then flat.
Connections
How dependent the outcome is on the wiring, and how easily that wiring could be reconstructed by a competitor.
Logic
How sensitive the outcome is to how the model is aimed, holding the model and the data fixed.

The legal-technology study in the following sections is one applied example of this general method.

Deep dive · Units

The first component, the unit: the model has become the baseline, commoditising the price of basic capability.

A base model is now the default unit of capability. It is the basic competence a product gets out of the box, before any wiring, data, or tuning. The quality of the baseline unit rose sharply, and the price for equivalent quality fell by roughly ten times a year. A capability that cost about sixty dollars per million tokens in 2021 costs a few cents today.

~10×
cheaper each year for the same quality, about a thousandfold over three years.
Exhibit 4 · Cost of equivalent model quality, 2021 to 2026
$60 $6 $0.60 $0.06 2021 2026 Year · Cost per million tokens, log scale $60 / M tokens ~$0.06 / M tokens
Source: a16z LLMflation analysis; public inference pricing.
Exhibit 5 · Market value, prior peak versus 2026 (indexed, peak = 100)
100 50 0 -72% Figma -78% Atlassian -60% Adobe Prior peak 2026
Source: public market data, indexed to prior peak.
Deep dive · Units

The companies repriced hardest were built closest to the baseline unit.

Figma's strong unit was a design-surface capability. Atlassian's was a workflow capability. Adobe's was a file-format and editing capability. Baseline generative models began producing and editing those artefacts directly. Adobe's forward multiple fell from roughly thirty times to twelve.

These were well-engineered companies with strong products and real customers. What they shared was proximity to the baseline unit, and that is what drove the repricing. A product whose advantage sits in data or connections is a separate case.

Deep dive · Units

A unit earns value back when it stops being generic.

The collapse in the baseline price does not make units worthless. It moves the value to units that a competitor cannot simply rent. A unit regains durable value in three main ways. Each turns the unit from a shared commodity into something specific to the product, and each is measurable.

The question for any product built on a model is simple: is its unit still the generic one, or has it become one of these? The answer sets how exposed the product is to the next model release.

01

Holds unique data in itself

A unit trained on data no one else has (proprietary weights, not just architecture) stops being generic.

02

Is specialised or fine-tuned

A model tuned for a narrow domain (tabular data, a legal corpus, a modality) can beat a larger general one on that domain, at lower cost.

03

Runs where the general model cannot

On-device, in a secure environment, or under latency and cost limits a frontier model cannot meet.

Deep dive · Data and knowledge

The second component, data and knowledge: a data advantage is real, but it runs out at a different point for every product.

Data helps until it stops helping. Adding examples improves a product quickly at first, then less, then barely at all. The point where the curve flattens is where the data advantage is effectively spent.

Where that point sits differs by product. A product with genuine compounding infrastructure (an email-deliverability system, where more volume keeps improving performance) keeps climbing. A product whose data advantage is shallow flattens within a handful of examples. The plateau location is specific to each product and has to be measured, not assumed.

Exhibit 6 · Score against data supplied, by product archetype
0.8 0.4 0.0 0 50 Examples or data supplied (indicative scale) · Product performance (composite score) Very deep Deep proprietary Moderate depth Close to baseline
Deep dive · Data and knowledge

In one product we measured, the data advantage was spent within five examples.

We took one contract-review product, the kind seen in a legal-technology deal, and measured how much its performance improved as we supplied more worked examples, scoring each result against answers expert lawyers had marked correct.

The first five examples moved the score by 0.286. The next forty-five moved it by 0.041. The advantage was front-loaded and then flat. Past a handful of examples, more data bought almost nothing.

+0.286
First five examples
+0.041
Next forty-five
Exhibit 7 · Score versus worked examples, Caliper Lab experiment
↔ Drag the point along the curve
0.8 0.4 0.0 0 5 50 The climb The flat Worked examples · Composite score
0.279
Composite score
0
Examples supplied
Composite score combines task accuracy, completeness, and expert agreement. Source: Caliper Lab benchmarks.
Exhibit 8 · Score as each layer is added, same contract-review product
0.4 0.2 0.0 0.284 Plain 0.267 +Retrieval 0.282 +Definitions The engineering layers barely moved the score
Source: Caliper Lab benchmarks.
Deep dive · Connections and logic

Holding the unit and data fixed, only how the model was aimed moved the score.

This is the third and fourth component read on the same product from the same study. With the unit and the available data held fixed, we rebuilt the work one layer at a time. Retrieval added nothing. Definitions added nothing. The parts that looked like the engineering moat moved the score barely at all.

The parts that looked like the engineering moat moved the score barely at all. What appears to be defensible plumbing often is not. Each layer has to be tested on its own before it can be credited with the advantage.

Case studies

The advantage sits in a different component each time.

Three products Caliper Lab has evaluated, anonymised here and read through the same four-component method. The same analysis places the advantage in a different layer each time, and rarely where the pitch would point.

Contract-review legal tool
A tool that reads contracts and extracts key terms, built on a rented general model.
Advantage in
Logic
The advantage is in how the model is aimed. The model is generic and the data runs out fast, so the only defensible part is the earned way the product directs the model at the task.
Units
Rented, generic
Data and knowledge
Shallow, spent in five examples
Connections
Incidental, rebuildable
Logic
Carries the advantage
Research and citation platform
A tool that surfaces and cites source material, built on a rented general model.
Advantage in
Data and knowledge
The advantage is in the data. The model is undifferentiated, but the proprietary corpus and citation graph are deep and hard for a competitor to reconstruct.
Units
Rented, undifferentiated
Data and knowledge
Deep, hard to rebuild
Connections
Standard, reproducible
Logic
Competent, not distinguishing
Forecasting and planning tool
A tool that forecasts on tabular and time-series data, built on a purpose-built model.
Advantage in
Units
The advantage is in the unit. The model is specialised for tabular data and beats any rented baseline, so the value sits in the model itself rather than the wrapper.
Units
Specialised, beats the baseline
Data and knowledge
Moderate, plateaus
Connections
Rebuildable
Logic
Standard
Source: Caliper Lab benchmarking.
How we do it

Every read is built on a proprietary base, not a single script.

Built on top of

Proprietary use-case reads

A library of completed reads across products and sectors, each mapping where a real asset's advantage sits.

Large knowledge base of experiments

Accumulated measurement runs across models, tasks, and configurations, so each new read starts from evidence, not from zero.

Pipeline of continuous runs and analysis

Standing infrastructure that re-runs evaluations as models change, tracking how fast a baseline is closing on a given product.

Where this applies

The same read informs both sides of an AI decision.

In each case the question is the same: what actually carries this product's performance, how durable is it, and what would a competitor have to reproduce.

For an investment decision
Assessing the risk of AI substitution to a target's product or moat.
Understanding a target's real AI and technological capabilities, not just its claims.
Tracking how model capabilities are moving, and what that means for a target's durability.
For an operating decision
Build versus buy, judged on how reproducible a product actually is.
Vendor selection, on measured task performance against your own ground truth.
Knowing where a current tool is exposed as models improve.
Caliper Lab

We assess what products and models can actually do.

Caliper Lab is an independent AI research and evaluations firm. We read the product itself, component by component, against independent ground truth.

Built by a team of AI researchers, scientists, machine learning practitioners, and domain leaders.

Independent

An outside read, with no stake in the products we measure.

Traceable

Every number traces back to an independently constructed task and a logged run.

Continuous

Continuous runs track how fast the competitive baseline is closing.

Get in touch

We are actively expanding our library of use cases and evaluations. If you have a product to assess, or a question about how we work, we would be glad to talk.

Dhruv Gulati  ·  dhruv@thecaliperlab.com  ·  thecaliperlab.com