---
title: Semantic Site Search for Publishers
description: Site search for publishers that matches meaning, not keywords — semantic search over your vectorized archive, with zero-result queries rescued and logged.
url: https://preview.artificialpoets.com/use-cases/intent-search/
site: Artificial Poets
type: page
date: 2026-08-09T19:46:31+00:00
modified: 2026-08-18T02:10:14+00:00
image: https://preview.artificialpoets.com/wp-content/uploads/a13s-cards/1501-social-10778d38.png
---
# Search That Understands the Question

**Measured on live deployments**

"Readers search my site and leave empty-handed — can search understand what they're actually asking?" It can, if it matches meaning instead of words. The Platform runs search over the same vectorized archive that drives its recommendations, so the query is matched against what every article is about and answered with your journalism.

[See it on my numbers](/request-a-demo/)

Horse & Rider · Ride TV · Equine Network · Equus Magazine · The Score · Equine Network Lockup

## The problem, in your numbers

A typed query is the clearest intent signal a reader ever gives you — first-party, consented, in their own words. Most search boxes answer it with string matching.

- **72%** Consumers who abandon a site when they can't find what they need — 53% go to Google, 36% to a competitor (Coveo, 2025)
- **10–15%** Industry-average share of site searches returning zero results; 20%+ on poorly tuned setups (Hello Retail, 2026)
- **56%** Sites with mediocre-or-worse on-site search UX in Baymard's 2026 benchmark
- **4% vs 19%** Click-through to the source from AI chatbot answers vs from search engines (Reuters Institute, 2026)

## The same query, matched two ways

Keyword search finds strings. Semantic search finds the articles about proofing temperature, starter health and hydration that answer the question — none of which contain the reader's words.

- **Keyword match** — The query "why won't my sourdough rise" is looked up against pages containing the exact words only. Zero results, and the reader leaves.
- **Semantic match** — The same query maps to its concept neighbors — proofing temperature, starter health, hydration — and is answered from your archive.

## The evidence, at its honest tier

- **The index answering search queries is the same vectorized archive measured on live deployments** — matching on meaning moved multi-page sessions +46% on enabled titles (10.3% → 15.1%) while comparison titles fell −16% (12.0% → 10.1%) over the same nine months
- **Largest single-title lift** — +19.1% in the first quarter, +4.0% at steady state — both published; plan on the second
- **Readers follow the matches voluntarily** — continuation past the fourth served article runs 70–82%, with a ~19-second per-article reading floor — meaning-matched content holds attention that keyword lists don't
- **What we do not claim:** — we publish no search-to-click conversion figure yet. The search funnel is instrumented from day one of a deployment and reported against its own baseline — not borrowed from someone else's benchmark

The index doing that matching is the one measured here: [two titles grew multi-page sessions 46% while the comparison group fell 16%](/customers/equine-network-rollout/).

**Introducing**

## Artificial Poets Platform

One engine behind every solution on this site. It learns your archive and your readers, then acts inside your CMS, your templates and your ad stack.

- It learns your archive: **Every story you have published, current again** — The engine understands each piece by what it is about, not when it ran or where it was filed. A feature from 2019 competes for the next slot on merit with one from this morning.
- It reads the visit: **What a reader wants, without asking** — Interest builds from what someone actually does in the session. No login, no third-party cookies, nothing leaving your domain. Useful on the second pageview, not the tenth visit.
- It chooses: **The right next read, not the popular one** — Someone comparing products and someone following a running story want different things. A most-read list gives both the same five links and serves neither.
- It serves: **There before the reader leaves** — The feed, the recommendations, the search answer and the signup ask all run on the same engine, in your templates and your ad stack. Any slot that arrives with them is yours to sell.
- It proves: **A lift you can defend, or we say so** — Every deployment runs beside titles that did not get it, plus a serving pause. That is how a result becomes a number you can take to a board instead of a vendor claim.

## FAQ

### What share of our searches currently comes back empty?

Most publishers cannot answer that, which is the first thing worth fixing: default site search rarely reports its own failures. Industry guidance puts zero-result rates around 10 to 15%, with good implementations under 5%. We would rather start by measuring yours than quote a lift number that belongs to someone else's archive.

### If a reader misspells a name or searches an acronym, does meaning-matching make it worse?

It can, which is why meaning-matching alone is not the whole answer. Embeddings are strong on synonyms and paraphrase and weak on exact tokens: proper nouns, tickers, acronyms. Those cases need literal matching alongside the semantic index, and evaluating them on your own queries is part of the baseline rather than a claim we make in advance.

### If search gets this good at answering, do readers stop opening articles?

No — results are links into your archive, not a generated answer. Nothing is synthesized that could misstate your journalism, and when the archive doesn't cover a query, the page says so instead of improvising. The zero-click question-and-answer session is already happening off your property — only 4% of news consumers click through from AI chatbot answers (Reuters Institute, 2026). Search that lands the reader in your reporting brings the question back onto a surface you own.

### How long after we publish is a story findable?

Indexing is continuous rather than a nightly rebuild, which matters most in the hours when a story is worth searching for. If a story is not findable in the window it matters, the search box is decoration, so treat this as something to verify on your own publishing rhythm during the baseline.

### Searchers convert 2–3× — is that your lift claim?

No, and be suspicious of any vendor who quotes it as one. The commonly cited multiple (4.63% vs 2.77%, Econsultancy data) is selection bias: readers who use search were already your most engaged visitors. Our measured figures come from a comparison design on live deployments, and a search deployment gets the same treatment — baseline first, frozen metrics, a comparison group.

### Only a minority of visitors touch the search box — why invest there?

Two reasons. Searchers are your highest-intent visitors, worth a disproportionate share of attention. And the query log is an asset beyond the searchers themselves: internal queries are the only consented, forward-looking first-party intent signal a publisher still owns after Google Zero. Zero-result queries map what readers asked for that you never wrote — see editorial intelligence.

### Is this a search box, or a chatbot that paraphrases our journalism back at readers?

A search box. It returns your articles, attributed and linked, and it does not generate prose over them. Studies of AI-generated news answers keep finding significant sourcing and accuracy problems; the point of matching on meaning here is to find your journalism, not to summarise it.

## Related

- [Nobody Measures Past Page Two](https://preview.artificialpoets.com/blog/nobody-measures-past-page-two/) — Publisher analytics has exactly two standard depth numbers: bounce rate, which tells you a reader did not reach a second page, and pages per session, an average that a handful of outliers can drag anywhere. Between and beyond them sits a report almost nobody runs — the full distribution of how deep sessions actually go — and it is where reader behaviour, loyalty economics and product effects become visible.
- [See What Readers Want Before You Commission It](https://preview.artificialpoets.com/use-cases/editorial-intelligence/) — "What should we publish next — and what does our audience want that we've never given them?"
- [Horse & Rider Grew Pages Per Session 19% Without Growing Traffic](https://preview.artificialpoets.com/customers/horse-and-rider/) — In the first quarter after enabling the Platform, one in six Horse & Rider readers went past the first page, up from one in ten. Three comparable titles moved less than a quarter of a point.
- [Session Depth Assessment](https://preview.artificialpoets.com/resources/session-depth-assessment/) — Two numbers from your analytics place your single-page rate and depth distribution against a continuously measured publisher panel.

## See how this works on your titles

Thirty minutes on your analytics. We tell you what share of your sessions stop at the first page, and what this would realistically move first for a setup like yours. You leave with the annotated read. No deck.

[See it on my numbers](/request-a-demo/)

How these numbers are made: [the measurement method](/solutions/measurement/). Figures from a multi-title publisher network measured continuously, August 2025 to May 2026. Last updated August 15, 2026.

## Questions this page answers

### What is semantic site search for publishers?

It is on-site search that matches meaning instead of words: the query is embedded as a concept and matched against what every article is about, so a reader who phrases something differently from your headline still finds the piece. It runs over the same vectorized archive that drives the recommendations.

### How does semantic search work on a publisher archive?

Every article in the library is embedded by subject rather than indexed by keyword, so a 2019 feature and this morning's post compete on merit for the same query. The reader's in-session interest shapes the ranking, and results are answered with your own journalism rather than sent elsewhere.

### How much does it cost to replace site search?

There is no public price list; pricing follows network size and formats. Search runs on the same vectorized archive as the recommendations, so it is an additional surface on one deployment rather than a separate product with its own integration and its own bill.

### Who is it for?

Publishers whose readers search and leave empty-handed. If your search returns nothing for queries your archive can answer, or ranks by date because that is all the index knows, the loss is invisible in most analytics: the reader asked a question you had already answered and did not get it.

### What are the alternatives to semantic site search?

The default CMS search matches words and usually orders by date. Hosted search products index your content well but match queries rather than meaning, and are priced as their own system. Here the index is the one already running your recommendations, which is also the index measured on live deployments.

### How do I get started?

Search is enabled on the same deployment as the rest of the engine: authorization, CMS access, infrastructure access, then four weeks of measurement-only baseline while the archive is indexed, then enablement on a named day. Your existing search results give you the before to judge it against.
