---
title: "Publisher Engagement Measurement: Audit-Ready Evidence"
description: "A publisher analytics platform measured the way your analyst would run it: four-week baselines, comparison groups, a published off-period, launch vs run rate."
url: https://preview.artificialpoets.com/solutions/measurement/
site: Artificial Poets
type: page
date: 2026-08-09T06:31:06+00:00
modified: 2026-08-18T02:10:14+00:00
image: https://preview.artificialpoets.com/wp-content/uploads/a13s-cards/1373-social-41daef74.png
---
# Numbers That Survive Your Analyst

**Measured on live deployments**

"How do I know any of this is real — and how do I prove it to my board?" Each number came from a design your analyst can rerun: baseline, a named enablement day, a comparison group of every unchanged title, casualties printed. The headline, +46% against −16% over nine months, holds when checked.

[See it on my numbers](/request-a-demo/) · [Get the claim audit checklist](/resources/vendor-claim-audit-checklist/)

Horse & Rider · Ride TV · Equine Network · Equus Magazine · The Score · Equine Network Lockup

## The problem, in your numbers

Distrust is the rational position. The fix is not a better chart: it is a checkable design.

- **38%** Publisher executives confident about the years ahead, against a −43% three-year traffic forecast (Reuters Institute)
- **78%** Marketing leaders who say martech failed its promised ROI (eClerx, May 2026)
- **44%** Marketers who question the reliability of their own incrementality results (EMARKETER/TransUnion)
- **94%** B2B buyers who fact-check AI outputs in a purchase; only 2% always trust them (TrustRadius 2026)

## The measurement stack

Five design layers between a dashboard delta and a number we publish, plus an observed off-period. This is how our own published numbers were made; every line is one your analyst can check.

- **Baseline** four weeks, collection and indexing only, zero serving
- **Enablement** a named day, on named titles; the series breaks cleanly
- **Comparison group** the entire eligible set of unchanged titles
- **Window sweep** the result holds when the start and end dates move
- **Segment cut** every device and channel, casualties printed
- **Off-period (observed)** serving paused four weeks, zero renders; the metric returned to baseline

## The evidence, at its honest tier

- **Multi-page sessions** — +46% on enabled titles (10.3% → 15.1%) vs −16% on comparison titles (12.0% → 10.1%) over nine months, on a multi-title publisher network we measure continuously
- **Launch vs run rate, side by side** — +19.1% first quarter → +4.0% steady state on the largest title; +12.2% → +6.9% on the second
- **The off-period is published** — across a four-week pause in serving — zero renders — the second title read +12.2% → −2.7% → +6.9%. The metric returned to baseline and recovered on resume
- **The segment cut is published with its casualty** — mobile +48% vs −23% comparison; direct fell on enabled titles too (−19%, vs −35% on comparison) — and we print it
- **No revenue figures and no traffic claims** — sessions and users were flat to slightly down. Depth is the measured result; revenue is your own RPM arithmetic

**Introducing**

## Artificial Poets Platform

One engine behind every solution on this site. It learns your archive and your readers, then acts inside your CMS, your templates and your ad stack.

- It learns your archive: **Every story you have published, current again** — The engine understands each piece by what it is about, not when it ran or where it was filed. A feature from 2019 competes for the next slot on merit with one from this morning.
- It reads the visit: **What a reader wants, without asking** — Interest builds from what someone actually does in the session. No login, no third-party cookies, nothing leaving your domain. Useful on the second pageview, not the tenth visit.
- It chooses: **The right next read, not the popular one** — Someone comparing products and someone following a running story want different things. A most-read list gives both the same five links and serves neither.
- It serves: **There before the reader leaves** — The feed, the recommendations, the search answer and the signup ask all run on the same engine, in your templates and your ad stack. Any slot that arrives with them is yours to sell.
- It proves: **A lift you can defend, or we say so** — Every deployment runs beside titles that did not get it, plus a serving pause. That is how a result becomes a number you can take to a board instead of a vendor claim.

## FAQ

### Aren't you grading your own homework?

The dashboard is ours; the record is yours. Serving events land in your analytics stack, metric definitions are frozen before the baseline starts, and the comparison group is the entire eligible set of unchanged titles — not a favorable selection. The published record also contains what self-graded homework never does: a launch figure that decays (+19.1% → +4.0%), a segment where enabled titles still fell, and a four-week off-period.

### How is the comparison group actually built?

It is the entire eligible set of unchanged titles in the same network, not a favorable selection. Cuts are within-segment, mobile against mobile, social against social, and the result must survive a window sweep and an observed off-period before we print it.

### Will the lift show in our own analytics, and what if your dashboard and ours disagree?

Serving events land in your analytics stack from day one, and metric definitions are frozen before the baseline starts, so the number is computed from your record, not ours. If dashboards disagree, yours is the one that counts: you were told to judge us from your own analytics.

### Engagement isn't revenue. Where is the dollar line?

We publish no revenue figures because we hold none of your revenue data — and a vendor's revenue claim is exactly the kind of number you should not accept. What we publish is measured depth and time: ~50 seconds median added reading per engaged session, ~30 seconds per served article. The revenue line is your arithmetic — your RPM against the measured depth.

### What happened after the novelty wore off?

It decayed, and we printed the decay: +19.1% in the first quarter fell to +4.0% at steady state on the largest title; the second title went +12.2% → +6.9%. The five-plus-article pool grew +221% and carries its caveat in the same sentence: under 2% of sessions reach five articles even after tripling. Plan on the steady-state numbers.

### What gets locked before a pilot starts, and what is the walk-away?

The metric, the success criteria, and the enablement day, all agreed before anything serves. The baseline month gives you a real before; the walk-away is deliberately cheap, a quarter and a JavaScript tag with nothing to unpick, and your record shows exactly what happened.

### Can our analysts interrogate the underlying record themselves?

That is the design: the record accumulates in your own analytics stack from the baseline onward, under definitions frozen in advance. Your analysts query it with their own tools, on their own access, and every published number here is one that survived that treatment.

## Related

- [The Vendor Claim Audit Checklist](https://preview.artificialpoets.com/resources/vendor-claim-audit-checklist/) — Fifteen checks in five groups — baseline, comparison, windows, mechanism, reversibility — for auditing any vendor’s engagement claim. Including ours.
- [How to Check an AI Vendor’s Numbers — Starting With Ours](https://preview.artificialpoets.com/blog/check-ai-vendor-numbers/) — Every AI vendor in publishing will show you a chart that goes up. The chart is real, the data is real, and the number can still be meaningless — because the decisions that matter happened before the chart was drawn: which window, which comparison, which metric, which segments were left out. Six checks expose those decisions. We run our own published numbers through each one, so you can see what passing and failing actually look like.
- [Publisher Session Depth Benchmark](https://preview.artificialpoets.com/resources/session-depth-benchmark/) — What normal looks like past page two — multi-page share bands, the full depth distribution, device and channel splits, and the rate depth erodes when nothing is done.
- [Analytics Health Check](https://preview.artificialpoets.com/resources/analytics-health-check/) — Ten one-query diagnostics that catch broken metrics before they decide anything. Each exists because it caught a real failure.

## See how this works on your titles

Thirty minutes on your analytics. We tell you what share of your sessions stop at the first page, and what this would realistically move first for a setup like yours. You leave with the annotated read. No deck.

[See it on my numbers](/request-a-demo/)

How these numbers are made: [the rollout case study](/customers/equine-network-rollout/) carries the full methodology. Figures from a multi-title publisher network measured continuously, August 2025 to May 2026. Last updated August 15, 2026.

## Questions this page answers

### How do you know an engagement lift is real?

By design, not by dashboard. A four-week measurement-only baseline, enablement on a named day, a comparison group of every unchanged title, a window-sensitivity sweep, a segment cut with the casualties printed, and an observed off-period: serving paused four weeks, the metric returned to baseline, then recovered on resume.

### How does the measurement design work?

Metric definitions are frozen before the baseline starts, and serving events land in your analytics stack, so the record is yours to recheck. The comparison group is the entire eligible set of unchanged titles, not a favorable selection, and the published record includes what self-graded numbers never do: a decaying launch figure and a segment that fell.

### What does it cost to find out if it works?

The measurement design is not an add-on; the four-week baseline and the agreed success criteria are how every deployment starts. There is no public price list, but the exit is priced in: if the number does not survive your analyst, what you spent is a quarter and a JavaScript tag, and your record shows exactly what happened.

### Who is this measurement approach for?

Analysts and boards who distrust vendor dashboards, rationally: 44% of marketers question their own incrementality results. Also the operators who need engagement evidence for ad buyers. Every number this site publishes was produced by this design, and the off-period is published alongside it.

### What are the alternatives to this kind of measurement?

Before-and-after deltas move with seasonality, news cycles, and traffic mix. Vendor dashboards grade their own homework. Whatever vendor you keep, including us, the Vendor Claim Audit Checklist is the fifteen checks we think you should run on every engagement claim, ours included.

### How do I get started?

Start with the baseline: four weeks of collection and indexing with zero serving, so the before exists in your own analytics. Freeze the metric definitions, name the enablement day, and let the unchanged titles be the comparison group. Eight weeks from authorization to a number your analyst can interrogate.
