---
title: How to Structure a Vendor Pilot You Can Actually Measure
description: Most vendor pilots end with a chart nobody trusts. The staged-rollout design — enable on a subset, hold comparable properties back, freeze metrics first — produces an answer instead of an argument. With a measured example.
url: https://preview.artificialpoets.com/blog/vendor-pilot-structure/
site: Artificial Poets
type: post
date: 2026-08-10T03:32:09+00:00
modified: 2026-08-14T18:30:48+00:00
categories:
  - Vendor Evaluation
---
# How to Structure a Vendor Pilot You Can Actually Measure

## Why default pilots can't answer the question

The question a pilot exists to answer is *"what did this change?"* — which requires knowing what would have happened without it. A pilot that deploys to everything at once destroys its own counterfactual on day one. From then on, every movement in the numbers is arguable: seasonality, a news cycle, an algorithm update, the redesign that shipped in week three.

The vendor will attribute the ups to the product; sceptics will attribute them to everything else; both are unfalsifiable, which is why pilot decisions so often revert to relationship quality and deck aesthetics. That is not an evaluation. It is a vibe with a spreadsheet attached.

## The design that produces an answer

Four decisions, all made before anything is enabled.

**1 · Split, don't blanket.** Enable on a subset of properties — or for single-site publishers, a subset of sections or templates — and leave genuinely comparable inventory untouched. The held-back group is the counterfactual: same audience, same seasonality, same news cycle, no product. The network we measure enabled on two titles of a multi-title portfolio and left the rest alone; when the enabled titles later moved +46% on multi-page share while the unchanged titles fell 16% over the same weeks, the design carried the whole argument.

**2 · Freeze the metric before the switch.** Name the primary metric, its exact definition, and the measurement windows in the pilot agreement itself. A success metric chosen after results exist is post-hoc selection wearing a suit — and it is the single most common way honest teams fool themselves. Freeze how it is counted, too: distinct articles per session, not raw pageviews, unless you enjoy [auto-load inflation](#).

**3 · Baseline first, enable second.** Insist on a measurement-only period before any feature turns on — four weeks is a reasonable floor. It anchors the intervention date unambiguously, captures the pre-period on identical instrumentation, and costs the vendor nothing but patience. A vendor who resists baselining is telling you which part of the pilot worries them. (The full argument gets its own article: [make every vendor wait four weeks](#).)

**4 · Bound the window by data integrity, not the calendar.** Decide in advance what invalidates a week — tracking gaps, migrations, retags on either group — and cap the analysis at the last clean date. We learned this one expensively: a comparison group whose tracking died mid-window [inflated a headline by roughly double](#) before anyone noticed.

## What the answer looks like when the design holds

With the four decisions in place, the final meeting changes character. Instead of duelling narratives, there is a table: enabled group versus held-back group, frozen metric, stated window. On the measured network it read +46% against −16% — and because the design existed, the follow-up questions had answers too. Segment cuts showed where the effect landed (mobile +48% against a −23% comparison decline). The launch-quarter figure could be separated from the steady state (+19.1% first quarter, +4.0% run rate on the largest title — plan on the second). And when serving later paused for four weeks on one title, the metric fell back to baseline and recovered on resume, which is the closest thing to proof a production environment offers.

None of that required a data-science team. It required four decisions made at signing time instead of at reporting time.

## The objections, quickly

**"Holding properties back costs us upside."** It defers upside on the held-back group by a few months, and in exchange converts the pilot from an anecdote into evidence. On the measured panel the held-back titles were *declining* — the "cost" of holding them back was learning that flat is what no-intervention actually looks like.

**"Our properties aren't comparable."** They don't need to be identical — they need to share seasonality and audience character well enough that their *trends* move together, which the baseline period itself will demonstrate or refute.

**"The vendor says their other case studies already prove it."** Other people's case studies prove what happened on other people's sites. A vendor confident in their effect should be eager to re-demonstrate it under your design — we are, because the design is where our numbers came from.

- The measured staged rollout: enabled titles **+46%** multi-page share, held-back titles **−16%**, same weeks
- Baseline floor: **4 weeks** of measurement-only before enablement
- Launch vs steady state on the largest title: **+19.1% → +4.0%** — freeze both windows in advance
- The blackout lesson: an unmonitored comparison group inflated a headline **~2×**

## FAQ

Baseline (4+ weeks) plus at least one full quarter enabled — long enough to see past the launch bump. Shorter pilots measure novelty.

Split by section or template rather than by property, and lean harder on the baseline and on off-period evidence. Weaker than a property split, still far stronger than before/after.

You, on your analytics, with the vendor's numbers reconciled against yours. A pilot measured only in the vendor's dashboard has one more conflict of interest than it needs.

## Structure yours on paper first

The Pilot Design Template walks the four decisions — split, metric freeze, baseline, window bounds — as a fill-in worksheet you can attach to the pilot agreement.

## See what your session depth looks like

We read your analytics with you for twenty minutes and tell you what share of your sessions stop at the first page. You leave with the annotated read, whether or not we ever talk again. 20 minutes, your analytics, no deck.

[Book my session-depth read](/request-a-demo/)
