---
title: Make Every Vendor Wait Four Weeks
description: A four-week measurement-only baseline before any vendor feature turns on is the cheapest credibility test in procurement — it anchors the intervention date, captures a clean pre-period, and reveals which vendors fear measurement.
url: https://preview.artificialpoets.com/blog/four-week-baseline/
site: Artificial Poets
type: post
date: 2026-08-10T03:32:09+00:00
modified: 2026-08-14T18:30:48+00:00
author: Matías Sanchez Moises
image: https://preview.artificialpoets.com/wp-content/uploads/a13s-cards/1593-social-f44fac24.png
categories:
  - Measurement
---
# Make Every Vendor Wait Four Weeks

## What the baseline actually buys you

**An unambiguous intervention date.** Every later claim — theirs or yours — hangs on knowing exactly when the treatment started. "We went live sometime in March, features ramped through April" is how attribution dies. Instrument on day zero, enable on a named day four weeks later, and the before/after boundary is a fact instead of a negotiation. On the deployments we run, that enablement date is the anchor for every published number.

**A pre-period on identical instrumentation.** The classic pilot mistake is comparing the vendor's post-launch measurements against your old analytics' pre-launch history — two different tools, definitions and failure modes, presented as a trend. Four weeks of the vendor's own collection before enablement gives you before-and-after on the *same* pipeline. The metric definition freezes at day zero, not at reporting time.

**A seasonality read on the split itself.** If the pilot holds comparable properties back — as [it should](#) — the baseline shows whether the enabled and held-back groups actually trend together before anything changes. If they diverge during a period when nothing is on, the comparison needed repair before it was used. Better to learn that in week three than in the results meeting.

**Time for the boring failures.** Tag conflicts, consent-mode gaps, bot inflation, double-firing — instrumentation problems surface in the first weeks of any integration. In a baseline they are caught and fixed while the data doesn't count. In an enable-everything-day-one pilot they are discovered later, inside the results, where fixing them and re-running is politically expensive.

## Why vendors resist, and what resistance tells you

A measurement-only month delays the vendor's "value moment," lengthens their sales cycle, and — uncomfortably for some — creates a clean record of what the site looked like without them. A vendor whose effect is real loses nothing but time. A vendor whose effect is a launch-week novelty bump, a seasonal coincidence, or a metric-definition change has everything to lose from a crisp baseline.

Which is why the request doubles as diligence. Ask for the four weeks and watch: enthusiastic agreement suggests a vendor who has been measured before and survived; "our other case studies already prove it" suggests the measurement lives somewhere you cannot check; "the algorithm needs to be live to learn" deserves the follow-up *"then how will we ever know what it added?"*

There is a legitimate version of the objection — some systems genuinely warm up on live data. The honest resolution is what we do ourselves: collect *and* prepare during the baseline (indexing, modelling, configuration), enable at the four-week mark, and accept that the model tunes further after enablement. Warm-up explains a ramp after the intervention date. It does not excuse the absence of one.

## What it looks like when the baseline exists

Every defensible number we publish traces back to this discipline. The enablement date is a named day, so launch-quarter and steady-state figures could be separated cleanly (+19.1% first quarter, +4.0% run rate on the largest title — both printed). The pre-period sits on the same instrumentation as the post-period, so definition drift is off the table. And when the effect was later challenged — by us — the baseline is what made the checks possible: window sweeps against a fixed anchor, comparison groups validated in their quiet weeks, an off-period read against a known floor.

Four weeks of patience at the start is what makes every subsequent claim auditable. There is no version of that trade a buyer should decline.

- Baseline floor: **4 weeks**, measurement-only, enablement on a named day
- What it anchors: the intervention date behind figures like **+19.1% → +4.0%** (launch vs run rate, both published)
- What it catches early: instrumentation faults, in the weeks the data doesn't count
- What it reveals: whether treatment and comparison groups trend together **before** anything is on

## FAQ

Two weeks is one bad news cycle. Four covers at least one full weekly rhythm repeated, catches month-boundary effects, and is short enough that no procurement timeline dies of it. More is better when seasonality is strong; less starts costing certainty.

It delays the *claim* of value by four weeks and brings forward the *proof* of value by forever. Deployment work — integration, indexing, configuration — proceeds during the baseline; ours does.

It is *a* baseline, on different instrumentation. Keep it as context, but the comparison that survives scrutiny is same-pipeline before/after. Your history also cannot validate the vendor's collection — the four weeks does both jobs at once.

## Put it in the agreement

The Pilot Design Template includes the baseline clause, the metric freeze and the window bounds as fill-in fields you can attach to any vendor contract.

## See what your session depth looks like

We read your analytics with you for twenty minutes and tell you what share of your sessions stop at the first page. You leave with the annotated read, whether or not we ever talk again. 20 minutes, your analytics, no deck.

[Book my session-depth read](/request-a-demo/)
