AEO Content Pipeline Data, v1

This page reports first-party measurements taken by iTech Valet from its own automated AEO content pipeline during 22 production runs on 28–29 August 2026. The runs researched and wrote articles on AEO and SEO topics. Every figure was computed once from the pipeline's live run database, at a recorded timestamp, and frozen. Nothing here is a survey, an estimate, or a third party's number.

Dataset summary

Summary of dataset AEO Content Pipeline Data version 1
TitleAEO Content Pipeline Data
Versionv1
PublisheriTech Valet
Measured byiTech Valet, from its own production systems
Subject areaAnswer Engine Optimization (AEO) and Search Engine Optimization (SEO)
Population22 production runs of one automated content pipeline, on one content cluster, for one client
Run window2026-08-28T19:36:52Z to 2026-08-29T00:13:21Z (4h 36m; the last run finished 00:15:53Z)
Measured on2026-08-29
Database read at2026-08-29T00:44:34Z
Frozen2026-08-29
Entries4
LicenceFree to quote and cite with attribution to iTech Valet and a link to this page

What this data is

iTech Valet runs an automated pipeline that researches and writes articles. For each article, the pipeline breaks the topic into fact-angles, searches for sources, filters the results, fetches pages, tries to extract a usable fact from each one, checks any figure it proposes to quote against the page it came from, and cites what survives. Every step of that is logged.

This page reports what those logs contain for one cluster of production work: 22 runs, on 28–29 August 2026, producing articles on AEO and SEO subjects. The figures were read directly from the pipeline's live run database on 2026-08-29 and frozen on the same day.

The measurements were taken by iTech Valet on iTech Valet's own system. That is the honest description of them and it cuts both ways: nothing here depends on anyone else's self-reporting, and nothing here has been independently audited.

Nothing on this page is derived when you load it. Each value was computed once and written down. Running the same query later, against a longer run history, will produce different numbers. When that happens the result is published as a new version that replaces this one, not as a correction appended to it. Two live numbers for one metric is the failure this practice exists to prevent.

Scope, and what it excludes

Every figure on this page is drawn from one population, described in full below. A figure quoted without it is being quoted about something this data did not measure.

The population

The population every figure on this page is drawn from
SourceThe pipeline's own run database (Cloudflare D1), tables runs and run_steps — read live and directly
ClientOne — iTech Valet's own content
ClusterOne — a 12-article AEO/SEO content cluster
Runs22 production runs — 10 finished, 12 refused or failed
Articles measured11 distinct articles — publish orders 2 through 12 of the cluster's 12. The first article's runs are excluded; see exclusion 2 below.
Fact-angles155
Run window2026-08-28T19:36:52Z to 2026-08-29T00:13:21Z (4h 36m; the last run finished 00:15:53Z)
Database read at2026-08-29T00:44:34Z
Measured on2026-08-29

The subject area is part of every number here

These runs researched AEO and SEO topics. In that vertical, a very large share of the indexable web is vendor marketing — agency blogs, tool landing pages, product comparison pages — which is exactly the category the pipeline's page-type gate exists to refuse.

So the refusal and extraction figures below are properties of this subject area meeting this pipeline. They are not a measurement of the web, and they are not a forecast for a clinic, a law firm, or a trade contractor. Every quotable sentence on this page names the subject area for that reason. A quotation that drops it is wrong even if the digits are right.

Three exclusions, each deliberate

  1. 488 proof-harness runs are excluded. They are real pipeline executions against invented topics under different quality settings, written by a test harness. They are 93% of the run table by row count. A query that forgets to scope to the real client returns a number that is mostly synthetic.
  2. 15 runs of the cluster's first article are excluded — which is why 11 articles are measured in a 12-article cluster. That article was re-run repeatedly while the pipeline itself was being built. Including those runs would weight one topic at 40% of the dataset and would mix pipeline-under-construction runs with production runs.
  3. A database backup was not used as the source. A backup taken earlier that day held only 9 of the real-client runs. Every figure on this page was read from the live database.

The measurements

Four entries. Each states its frozen value, what it counts, the sample it was drawn from, and the date range it covers. The block marked Quotable under each entry is the sentence iTech Valet asks to be quoted: it carries its own scope, so quoting it verbatim cannot strip the qualifier off the number.

1. The evidence-acquisition funnel

How much of what a search returns actually survives to be cited. Five stages, counted across the whole population.

Evidence acquisition, 22 production runs, AEO/SEO subject area
#StageCount
1Search results returned785
2Passed the pre-fetch filter681
3Pages actually fetched503
4Verified by the extractor201
5Cited in an article76

What was lost at each stage, and why

Every gap is accounted for.

Composition of each funnel gap
GapCountComposition
1 → 2 104 Duplicate URLs, and results refused by the domain filter stack.
2 → 3 178 72 left unread because the run's 30-fetch budget was exhausted; 106 left unread because no unread candidate could still outscore the best source already verified.
3 → 4 302 Nothing extractable 167 · blocked or challenged 62 · extractor error 40 · fetch error 30 · redirect 2 · paywalled 1.
4 → 5 125 110 refused by the quality gates — page type 103, figure not present in the source 5, per-domain cap 2 — and 15 eligible but beaten by a better candidate for the same angle.

The page-type refusal, stated with its denominator

103 refusals per 201 verified pages.

The largest single reason a verified page was not cited is that it was the wrong kind of page — vendor marketing rather than a source of fact. The gate only ever sees a page that was fetched and verified, so 201 is the only denominator this count belongs to. Expressing it against pages fetched, or against search results, would put it against a population the gate never judged.

Sample size
21 runs (one of the 22 died at the gate before any search), 155 fact-angles, 785 search results
Date range
2026-08-28T19:36:52Z to 2026-08-29T00:13:21Z
Measured on
2026-08-29

Quotable

Across 22 production runs of a 12-article AEO/SEO content cluster, the pipeline's searches returned 785 results; 681 passed its pre-fetch filter, 503 were fetched, 201 were verified by an extractor, and 76 were ultimately cited — roughly one citation for every ten search results returned.

2. Nothing-extracted rate

167 of 503 pages fetched — 33.2%.

A page that was fetched successfully, read by the extractor, and yielded no usable fact at all. Not blocked, not an error, not the wrong kind of page — a page that was read and had nothing in it worth quoting.

Sample size
503 fetched pages across 21 runs
Date range
2026-08-28T19:36:52Z to 2026-08-29T00:13:21Z
Measured on
2026-08-29

Quotable

In a measured run of an AEO/SEO content pipeline, a third of the web pages it fetched — 167 of 503 — were read in full and yielded no usable fact at all.

3. Blocked-or-challenged rate

62 of 503 fetch attempts — 12.3%.

The label is “blocked or challenged”, and it is deliberately not “firewall rate”. The underlying counter absorbs 403s, 429s and bot-challenge interstitials alike. The pipeline cannot tell a firewall from a rate limit from a challenge page, and a label that named one mechanism would be claiming a distinction the measurement never made.

No per-domain table is published, and the data is the reason. Blocking is not a stable property of a publisher in this dataset: the single most-cited domain in the wider run history was also blocked on other attempts. A per-domain table would read as “these publishers block us” when what was measured is “this share of attempts was blocked”.

Sample size
503 fetch attempts across 21 runs — 57 distinct URLs on 26 distinct registrable domains after folding the www. prefix
Date range
2026-08-28T19:36:52Z to 2026-08-29T00:13:21Z
Measured on
2026-08-29

Quotable

In a measured run of an AEO/SEO content pipeline, 12.3% of page fetch attempts — 62 of 503 — were blocked or challenged before the page could be read.

4. Hallucination interceptions

5 catches.

The composer proposed a citation carrying a numeric figure that was not present in the source page it cited. Code compared the figure against the fetched text, found it absent, and refused the citation before it could reach a draft.

The sample is small and the smallness is part of the figure. Five events across 201 verified pages in 21 runs. This is a count, not a rate, and it is published as a count. There is not enough here to support a claim about what share of AI-proposed citations misquote their source, and this dataset does not make one.

Sample size
5 events, from 201 verified pages across 21 runs
Date range
2026-08-28T19:36:52Z to 2026-08-29T00:13:21Z
Measured on
2026-08-29

Quotable

In 22 production runs of an AEO/SEO content pipeline, an automated check caught 5 occasions on which the drafting model attached a numeric figure to a source page that did not contain it, and refused the citation before it reached a draft.

How this was measured

The pipeline writes a step record for every run. Each record carries a JSON summary naming, per fact-angle, how many results a search found, how many passed the pre-fetch filter, how many pages were examined, how many were verified, and which source won. It also carries the reason every rejected, disqualified or unfetched candidate was set aside.

All figures were read from those records directly, in the live database, on 2026-08-29 at 00:44:34Z. No figure was estimated, interpolated, or carried over from an earlier count.

Source of each measurement
EntryHow it was counted
1. Funnel Stage counts summed across each run's per-angle outcome records — results found, results passed, pages examined, pages verified, and one count per outcome that produced a winning source. Where a run re-ran its angle selection, the later record supersedes the earlier one for that run, because its summary is cumulative and restates the whole run. Gap compositions come from the recorded reason on each rejected, disqualified and unfetched candidate.
2. Nothing-extracted Count of rejected candidates whose recorded reason was nothing_extracted, over the sum of pages examined. This gap is an exact partition — 302 rejections against a 302-page gap — so numerator and denominator are drawn from the same closed set.
3. Blocked or challenged Count of rejected candidates whose recorded reason was the block-or-challenge outcome, over the sum of pages examined. Host names were lower-cased and the www. prefix folded before any distinct count.
4. Hallucination interceptions Count of disqualified citations whose recorded reason was that the figure was not present in the source.

One reconciliation note, recorded rather than smoothed over

Funnel stages 3→4 and 4→5 reconcile exactly: the rejection and disqualification logs partition those gaps to the unit.

Stage 1→2 does not reconcile against the rejection log, and the reason is understood. That log is an event log, not a partition. When a second search resurfaces a URL the first pass already refused, that one URL is logged twice — so the log records 151 events against a stage gap of 104. The stage counters are authoritative and the rejection log is evidence, not arithmetic. The funnel above is built from the counters throughout.

Measured but not published, and why

Three figures were computed and then held back. They are named here so that their absence reads as a decision rather than an oversight.

Figures measured but not published here
FigureWhy it is not published
Wall-clock per finished article Measured, frozen, and held out of the public dataset by decision. Nothing is wrong with the figure — it is not withdrawn, corrected, or in doubt. It remains in iTech Valet's internal data map and is simply not published here. It appeared in the first version of this page on 2026-08-29 and was removed the same day; see the changelog.
Sourcing yield (hard claims per run) The pass mark moved inside the measurement window. Early runs read one evidence threshold and later runs read a lower one — and the refusal records show at least two of the early readings were a fallback default applied because a configuration cell was empty while it was being edited, not a setting anyone chose. A yield-against-threshold figure would pool two thresholds, one of which was an accident.
Cost per article A known defect in how cost was rolled up. The per-run cost column is written when a run reaches its terminal state and was not re-summed when a step ran afterwards, so any run with post-completion work under-reports its own cost. The defect is fixed in the code, but the fix is not retro-fitted to this frozen window, so no cost figure appears anywhere in v1.

Also unavailable, for reasons that are not judgement calls

  • Any starvation-rate figure. The telemetry that would measure it was built late and exists on only 2 runs in the entire history.
  • Any fleet or uptime figure. The tables that would carry it are empty.
  • Anything before 2026-08-27. The run registry begins there. No article produced before that date exists in the database at all, so no earlier comparison can be made.

How to cite this page

These figures are free to quote and cite with attribution. Where a measurement has a Quotable sentence above, quoting that sentence verbatim is preferred: it carries its own scope, so the number cannot arrive somewhere without the qualifier that makes it true.

iTech Valet. AEO Content Pipeline Data, v1. Published 2026-08-29. https://itechvalet.com/research/aeo-pipeline-data/

Changelog

— wall-clock removed from the published dataset

The fifth measurement — wall-clock time per finished article — was removed from this page later the same day. It is not withdrawn and not corrected: the figure is measured, frozen and unchanged in iTech Valet's internal data map, and is held out of the public dataset by decision. The four measurements above are unaltered, and no figure in any of them moved.

— v1 published

First publication. Five measurements frozen from 22 production runs recorded 2026-08-28T19:36:52Z to 2026-08-29T00:13:21Z, read from the live run database at 2026-08-29T00:44:34Z: the evidence-acquisition funnel, the nothing-extracted rate, the blocked-or-challenged rate, the hallucination-interception count, and wall-clock time per finished article.

Cost per article and sourcing yield were computed and deliberately withheld; both reasons are stated above.

A future measurement of the same pipeline is published as v2 and replaces these values on this page rather than appending to them. This changelog records the replacement, and the figures above are always the current frozen version.

Back to top
Processing...