opuscript.com/research
What do AI assistants say when someone asks who to hire?
Twenty Ottawa businesses across ten sectors, put to three AI assistants on an identical six-question instrument. Every answer captured verbatim and dated. The method and the predictions are published here before the data was collected.
Pre-registered · data collection in progress · results pendingWhy the method is published before the findings
A hypothesis written before the data is evidence. A hypothesis written afterwards is narration. Publishing the predictions first means they cannot be quietly adjusted to match whatever turns up — and it means you can check whether the eventual findings answer the questions that were actually asked.
This is not standard practice in marketing research, which is precisely why it is worth doing. Most published claims about AI visibility are assertions with a number attached and no way to tell whether the number was found or chosen.
The central hypothesis
Presence in AI answers is driven by third-party source density, not by website quality.
If that holds, it inverts the standard advice given to small businesses. It would mean the money spent redesigning a website buys very little in the answer layer, and the money spent getting listed, reviewed and written about buys nearly all of it.
The sample is built to test it directly. Within each of ten sectors, two real, established businesses are paired — one with high review volume, one with low. Both have websites. Both are trading. If the higher-density business is named and the lower-density one is not, sector after sector, the hypothesis survives.
The sample
Ten sectors, twenty businesses, paired within sector on Google review count as a proxy for third-party source density. Counts recorded 13 August 2026 and frozen as the independent variable. No business is named here, or in the published findings.
| Sector | Higher density | Lower density | Ratio |
|---|---|---|---|
| Accounting / bookkeeping | 414 | 34 | 12× |
| Home renovation | 125 | 25 | 5× |
| Physiotherapy | 474 | 6 | 79× |
| Wedding photography | 182 | 67 | 2.7× |
| HVAC | 1,465 | 125 | 12× |
| Picture framing | 60 | 39 | 1.5× |
| Landscape design | 240 | 64 | 3.8× |
| Counselling / psychotherapy | 55 | 12 | 4.6× |
| IT managed services | 76 | 1 | 76× |
| Custom furniture | 56 | 2 | 28× |
Picture framing, at 1.5×, is deliberately kept as a control. If presence differs sharply between two businesses with near-identical review counts, source density is not the whole story — and that would be worth knowing too.
One thing the sample revealed before any data was collected. An initial search for physiotherapy clinics returned four branches of a single chain, each with more than three hundred reviews. One operator had effectively saturated the category listing. A neighbourhood-level search found independent practices instead, and that sector now carries the widest density gap in the study.
The instrument
Six questions, identical in structure for every sector, filled from two variables only so results can be pooled. Questions 1 to 4 are asked once per sector and both paired businesses are scored against the same answers; questions 5 and 6 are asked for each business.
| Q | Type | Wording |
|---|---|---|
| 1 | Category | Who does {category} in Ottawa? |
| 2 | Category, ranked | Best {category} in Ottawa |
| 3 | Problem-first | {problem} Who should I call in Ottawa? |
| 4 | Forced recommendation | Recommend one {category} in Ottawa and explain why. |
| 5 | Qualification | Is {business} in Ottawa any good? |
| 6 | Specific fact | Where is {business} located and what do they charge? |
Question 4 carries the most weight. Forcing a single recommendation removes the hedge. An assistant listing six businesses tells you little; an assistant naming one tells you who actually wins the query — and it is the closest analogue to what a real customer does with the answer they receive.
Assistants: ChatGPT, Gemini and Perplexity. Three rather than four, chosen for consumer share on local-recommendation queries. Total captures: 240.
Pre-registered hypotheses
Six predictions, each with the condition that would falsify it. All six will be reported as confirmed or falsified. A study that confirms everything it predicted should be read with suspicion.
Within matched pairs, the higher-density business is named more often than the lower-density one in at least seven of ten sectors.
Falsified if the split is five-five or worse, or if website quality better explains the variance on inspection.
Median presence across all twenty businesses on questions 1 to 4 is below 25%.
Falsified if median presence exceeds 40%.
Presence on question 3 is at least fifteen percentage points lower than on question 1, because problem-first phrasing requires a connection nobody has written down.
Falsified if question 3 presence matches or exceeds question 1.
In at least 40% of forced-recommendation cells, assistants decline to name a specific local business — deflecting to a directory, a national brand, or advice to search locally and check reviews.
Falsified if a specific local business is named in more than 75% of those cells.
Questions 5 and 6 produce at least one error per business on average, with staleness and fabrication dominant — assistants invent pricing rather than decline to state it.
Falsified if fewer than ten total errors appear across forty named-business cells.
The three assistants name overlapping but substantially different businesses. Fewer than half of named businesses are named by all three on the same question.
Falsified if agreement on named businesses exceeds 70%.
How the twenty businesses are treated
The design deliberately rejects the anonymise-and-tease approach — publishing that one of twenty businesses is invisible and inviting each to wonder whether it is them. Ottawa is small enough that sector plus finding is often identifying, and a vague worry is not useful information.
- Aggregate and sector level only. No business is named in the published findings.
- Every business receives its own results, free and unrequested — its own answers, verbatim, with what was wrong and what to correct. The same artefact the audit service sells, given away at study scale.
- No business is contacted before data collection. Telling them first would contaminate the data. Telling them afterwards, with their results in hand, is the point.
- No negative verbatim tied to an identifiable business is published. If a quote cannot be anonymised, it is not used.
- Removal on request. Any business asking to be excluded is excluded, and the reduced sample size is reported.
What the study will not be able to claim
n = 20. Ten matched pairs supports a directional finding, not a statistical one. No p-value will be attached to it.
Review count is a proxy for third-party source density, not a measurement of it. It correlates with press, directory listings and mentions; it does not equal them.
Correlation only. Even if H1 holds perfectly, the study cannot show that acquiring reviews causes AI presence. Both may follow from a third factor — age, size, or simply how much other people write about a business.
One city, one month, three assistants, one auditor. Answer engines are non-deterministic. A re-run would produce different specifics; the pattern is the claim, not the numbers.
Selection via Google Places biases toward businesses maintaining a Google profile at all — itself a source-density signal. This narrows the range and makes H1 harder to confirm, not easier.
Citing this study
McGrath, K. M. F. (2026). The Ottawa Answer-Engine Study: protocol and
pre-registered hypotheses. OPUSCRIPT, CIONAOD Inc.
https://opuscript.com/research/
The protocol, the instrument and the aggregate data are published under a Creative Commons Attribution licence. Reuse, replicate in another city, or argue with the findings — all of that is welcome and none of it requires permission. Replication in a second city would be more useful than agreement.
Be told when the results publish
One message when the findings are out, including the aggregate data file. No newsletter.