---
title: How to Check an AI Vendor’s Numbers — Starting With Ours
description: Every AI vendor shows publishers a chart that goes up. Six checks separate measured results from manufactured ones — windows, comparison groups, the off-switch test — worked openly against our own published numbers.
url: https://preview.artificialpoets.com/blog/check-ai-vendor-numbers/
site: Artificial Poets
type: post
date: 2026-08-10T03:32:09+00:00
modified: 2026-08-14T18:30:48+00:00
categories:
  - Vendor Evaluation
---
# How to Check an AI Vendor’s Numbers — Starting With Ours

Fewer than four in ten news executives say they are confident about the industry they work in.[[1]](#ref-1) That scepticism is earned, and the correct response to it is not louder claims. It is checkable ones.

## Check 1 — Ask for the window, and the run rate next to the peak

Launch quarters are real and they are always the biggest. An effect measured over its first twelve weeks will usually be a multiple of its steady state, because novelty, tuning and attention all peak early. A vendor who quotes the launch window without the run rate is not lying — they are choosing which true number you see.

**Ours:** the title with our largest published lift ran **+19.1%** in distinct pages per session over the first quarter after enablement — and **+4.0%** at steady state six months later. Both figures are in the published case study, in the same table. The +4.0% is the one to build a business case on, and we say so in the document.

## Check 2 — Ask what it was compared against

A before/after on one site proves nothing by itself; seasonality, news cycles and platform changes all move the after. The minimum honest answer is a comparison group. The stronger answer is that the comparison group is the *entire eligible set* — every similar property available — rather than a selection that happens to flatter the result.

**Ours:** the network we measure enabled our Platform on two titles and left the rest unchanged. Across the first quarter, distinct pages per session on three comparison titles moved **−0.2%, −0.0% and −0.1%** while the enabled titles moved +19.1% and +12.2%. The comparison set is every title in the network that reported continuously and was never enabled — the selection rule is in the methodology note of each case study.

## Check 3 — Ask what happens when it's switched off

A comparison group is still an argument about different sites. The strongest evidence any vendor can hold is the same site, same audience, same season, with the product removed — and engagement returning to baseline, then recovering when it resumes.

**Ours:** for four weeks in spring 2026 our Platform did not serve on one measured title. Distinct pages per session went from **+12.2%** above baseline to **−2.7%** — statistically back to where the site started — and recovered to **+6.9%** within a fortnight of serving resuming. Two independent signals date the pause identically. We did not plan it, and it is the best evidence we have.

## Check 4 — Ask for the metric's exact definition

"Engagement" is not a metric. Neither, on a modern site, is a raw pageview count — auto-loaded content, refreshes and instrumentation changes all move it without any reader doing anything. The definition needs to be specific enough that your own analyst could recompute it.

**Ours:** we publish *distinct article paths per session* — unique articles a session actually reached — precisely because raw pageviews failed our own robustness testing: under different measurement windows the raw-pageview "lift" swung from +32 points to −4, so we rejected it. The metric that survived moved less than a third of a point under the same test. The full sweep is published, and it is worth [reading before you evaluate anyone](#), including us.

## Check 5 — Ask which segments moved, and which didn't

An aggregate lift can be one strong segment carrying six flat ones — or worse, a mix-shift artifact with no real change anywhere. The segment cut is where manufactured numbers go to die, which is why you rarely see one in a sales deck.

**Ours:** six of seven segments improved on the enabled titles — mobile **+48%** against a comparison-group mobile decline of 23%, social arrivals roughly doubled. The seventh, direct traffic, **declined 19%**. It declined 35% on the comparison titles, so we did not cause it — but we did not fix it either, and the red bar is printed in the case study, because a chart with one of your own bars pointing the wrong way is read as a measurement rather than a claim.

## Check 6 — Ask what they will not claim

Every real intervention has a boundary. A vendor who cannot name theirs has not measured it — or has and prefers you didn't ask. The answer also tells you what kind of partner you are buying: the one who manages your expectations before the contract, or after.

**Ours, on the record:** traffic did not grow. Sessions and users were flat to slightly down on the enabled titles across every window we measured — our product deepens sessions; it does not acquire them. We claim no search-ranking effect. And we publish no revenue figures, because we do not have our customers' revenue data: any revenue arithmetic in our materials is the reader's own RPM applied to measured pageview changes, labelled as exactly that.

## Run it against us

The six checks are the audit. Run them on our published material — four case studies and a white paper, every figure carrying its window, its comparison and its methodology note — and then run them on whoever else is in the deck beside us.

If both vendors survive, you have two good options. If only one volunteers the material before you ask, that is also an answer.

- Launch quarter vs run rate on our largest published lift: **+19.1% → +4.0%**, both printed
- Comparison group: three titles at **−0.2% / −0.0% / −0.1%** while enabled titles moved +19.1% and +12.2%
- The off-switch: **+12.2% → −2.7% → +6.9%** across a four-week serving pause and recovery
- Rejected our own biggest metric after a **+32pp → −4pp** window sweep
- Published segment cut includes the one that fell: direct, **−19%**

## FAQ

It is unusual, not unfair. A vendor sitting on twelve months of live measurement has the data to answer all six checks. If the answers exist and are not offered, that is a pricing signal.

Then the honest sale is a structured pilot that creates the evidence: enable on a subset, hold comparable properties back, freeze the metric definitions first. A new vendor proposing that design is more credible than an established one waving a blended chart.

Checks 1, 2, 5 and 6 are questions in a meeting — they need no tooling, only the willingness to sit through the pause after you ask. Checks 3 and 4 take an analyst an afternoon with your own analytics.

## The fifteen-point version, for the procurement meeting

The Vendor Claim Audit Checklist expands these six checks into fifteen — what to ask, the red flag, and what a pass looks like, printable for the room.

## References

1. Reuters Institute for the Study of Journalism, “Journalism, Media and Technology Trends and Predictions 2026.” [reutersinstitute.politics.ox.ac.uk](https://reutersinstitute.politics.ox.ac.uk/journalism-media-and-technology-trends-and-predictions-2026)

## See what your session depth looks like

We read your analytics with you for twenty minutes and tell you what share of your sessions stop at the first page. You leave with the annotated read, whether or not we ever talk again. 20 minutes, your analytics, no deck.

[Book my session-depth read](/request-a-demo/)
