AI visibility audit methodology
An AI visibility audit is credible only when another analyst can inspect the ruler. This is the ruler 10xSearch uses for its new native audits, including the fixed panel, tracked fields, raw evidence, verdict rules, and limitations.
A credible AI visibility audit uses a frozen, versioned prompt panel and retains enough evidence for another analyst to independently reproduce every headline. The 10xSearch growth-audit-50.v1 panel contains 50 fixed prompts: 14 national category prompts, 12 luxury real estate specialization prompts, 12 competitive vendor prompts, and 12 problem and solution prompts. Geography is disabled unless the business intentionally wants local relevance. Each engine check stores the prompt, engine, provider model, execution scope, run date, timestamp, raw answer, cited URLs, mention and recommendation verdicts, exposed position, sentiment, provider request ID when available, error state, and SHA-256 capture hash. The first 10xSearch run produced 100 captures across Gemini and Claude with zero errors, 39 responses containing cited URLs, and zero mentions or recommendations of 10xSearch. Recognition, citation, recommendation, ranking, traffic, leads, and revenue remain separate measures.
Versioned as growth-audit-50.v1
Gemini and Claude, named with their provider models
One retained observation for every prompt-engine check
Tracked separately from mentions and recommendations
The fixed prompt panel
The prompt list is deterministic and versioned. A model does not invent the questions during the audit. That makes repeated runs comparable and stops one favorable or accidental prompt from becoming the benchmark.
- National category discovery: 14 prompts
- Luxury real estate specialization: 12 prompts
- Competitive and vendor comparisons: 12 prompts
- Problem and solution research: 12 prompts
- Intentional geography: zero for 10xSearch; four prompts replace problem prompts only when local relevance is deliberate
Fields retained for every check
Each prompt-engine result retains prompt ID, category, prompt text, engine, model, execution location, market scope, run date, timestamp, mention, recommendation, exposed position, sentiment, cited URLs, raw answer, provider request ID when available, error, and a SHA-256 hash of the capture.
How recommendation rate is calculated
Recommendation rate equals successful responses that pass the deterministic recommendation verdict divided by all successful prompt-engine responses. Errors stay visible and do not enter the successful denominator. Recognition, citation, position, and sentiment remain separate fields.
A self-audit that changed this site
On August 13 we crawled our own public site and found 402 sitemap URLs, 49 template pages visibly publishing an internal outline label, 363 non-blog URLs stamped with the current date, and only 16 of 402 sitemap URLs with external source links beyond shared navigation. Those findings triggered this consolidation, noindex, durable last-modified, and authority-page rebuild.
The first published 10xSearch baseline
The August 13 production run submitted all 50 fixed prompts to Gemini and Claude, yielding 100 raw hashed captures with zero errors. Thirty-nine responses included one or more cited URLs, but none mentioned or recommended 10xSearch. That is the baseline to beat, not a favorable one-off query.
How prompts enter and leave a scored panel
Prompt research and scored measurement are separate activities. Candidate questions can come from customer interviews, sales calls, search queries, competitor comparisons, product questions, and exploratory model sessions. Before scoring, each candidate is assigned to an intent family, checked for duplicate meaning, reviewed for accidental brand or location bias, and written so one engine run does not depend on prior conversation. The approved list is frozen with a panel ID. A prompt can be retired when the commercial decision disappears, wording proves ambiguous, or the business changes scope, but that change creates a new panel version. Historical results remain attached to the version that produced them. A bridge analysis may run the old and new panels in parallel, but the report must not blend their denominators silently. This discipline prevents a favorable question from entering after the baseline or an unfavorable question from vanishing before the comparison. It also makes the panel a durable operating asset rather than a one-time demonstration.
- Maintain a research pool outside the scored panel.
- Assign every prompt a stable ID, intent family, text, market scope, and inclusion rationale.
- Review branded wording, geographic assumptions, and semantic duplicates before freezing the list.
- Create a new version for additions, removals, wording changes, or category changes.
- Use bridge runs when a business needs continuity across materially different panel versions.
The raw-capture and provenance contract
Every prompt-engine observation needs enough provenance to distinguish what the provider returned from what the audit later inferred. The request record identifies the prompt, engine, provider model, execution scope, timestamp, and provider request ID when one exists. The response record retains the unedited answer text, cited URLs, completion or error state, and a cryptographic hash of the capture. Deterministic processing adds mention, recommendation, position, sentiment, and citation-host fields without overwriting the raw response. The run record freezes the panel version, subject identity, aliases, market scope, start and completion times, expected observation count, and artifact hash. Public reports may redact protected provider metadata or client-confidential content, but the durable database receipt must keep the internal evidence needed for audit and retry. A screenshot is useful visual corroboration, yet it cannot replace text that can be hashed, searched, reconciled, and compared across interfaces.
- Store raw provider output before applying verdict logic.
- Keep request identity, response identity, and derived verdicts in separate fields.
- Hash the retained capture and the complete public artifact.
- Record expected and actual observation counts so missing cells cannot disappear.
- Publish redaction and omission rules when the public artifact is narrower than the durable receipt.
Definitions for recognition, recommendation, citation, position, and sentiment
Recognition means the answer prose identifies the audited subject or an approved alias. A URL-only match does not become recognition unless the methodology explicitly defines and labels that broader rule. Recommendation requires language that presents the subject as a suitable choice for the prompt's decision, not merely as an example or source. Citation records the URLs exposed by the engine, whether or not those URLs belong to the subject. Position describes where the first qualifying recognition or recommendation appears in the ordered answer. Sentiment classifies the immediate context as positive, neutral, mixed, or negative under a published rule. These dimensions can diverge. An answer may cite a 10xSearch page but recommend another provider, mention 10xSearch negatively, or name the company after several competitors. The report should preserve each outcome separately. Deterministic rules make large panels repeatable; sampled human review checks whether those rules still match the intended commercial meaning.
- Publish approved subject names and aliases with the run.
- Do not count a source URL as prose recognition by default.
- Require recommendation language to match the commercial intent of the prompt.
- Capture the first qualifying position and the surrounding sentiment.
- Sample positive and negative verdicts for human quality review after logic changes.
How errors, retries, and missing observations are handled
A completed run is not merely a report page with a score. It has an expected observation count equal to prompts multiplied by configured engines, and every expected cell reaches a terminal success or terminal error state. Timeouts, provider refusals, rate limits, malformed responses, and parser failures remain visible. A retry creates a traceable attempt and does not erase the earlier failure. If the provider returns a valid answer without citations, that is a successful uncited observation, not an error. If an engine is unavailable for a material share of the panel, the report should disclose the coverage gap and avoid comparing the partial result to a complete baseline as if the denominators matched. Publication waits for the durable run and observation records, not only for an in-memory task to finish. This distinction matters because a clean-looking percentage can hide missing rows. Reconciliation checks expected cells, unique prompt-engine keys, raw captures, hashes, and artifact contents before the run becomes a public baseline.
- Compute the expected prompt-engine matrix before requests begin.
- Persist every attempt, terminal error, and successful uncited answer.
- Prevent duplicate attempts from inflating the scored denominator.
- Disclose material provider coverage gaps beside any comparison.
- Reconcile database rows, hashes, and the rendered artifact before publication.
How before-and-after comparisons remain compatible
A defensible comparison keeps the subject identity, panel version, aliases, verdict contract, engine, model family where practical, and execution scope visible for both periods. Some provider drift is unavoidable because models change, so the report should name that change instead of presenting the second run as a controlled laboratory experiment. The before state is frozen and retained. The after state runs only after material work is live and accessible. Results compare counts and rates using the same successful-denominator rule, while errors and engine coverage appear alongside the headline. Page changes, new external sources, profile corrections, and publication dates form an intervention ledger; they show what happened between runs without proving that any one change caused a model response. A bridge analysis is required when the prompt panel or verdict logic changes. Traffic and pipeline outcomes use their own comparable windows and attribution definitions. This approach supports operational learning while avoiding stronger causal claims than the evidence allows.
- Freeze the baseline artifact and never regenerate it in place.
- Record material site, source, profile, and measurement changes between runs.
- Compare compatible prompt, engine, scope, and verdict dimensions explicitly.
- Name provider-model drift and denominator differences instead of hiding them.
- Treat observed movement as evidence for prioritization, not automatic proof of causation.
Public evidence, client confidentiality, and retention
Transparency does not authorize disclosure of client-confidential information. The measurement contract begins by defining which subject names, prompts, answers, citations, screenshots, performance metrics, and customer statements may be public. A public evidence asset can expose the prompt panel, dated verdicts, cited URLs, hashes, and redacted captures while the service-role receipt retains protected details. Redaction must be consistent and disclosed; it cannot remove unfavorable observations or alter the scored denominator. Customer outcomes are published only at the level supported by the source and permission. A live-site credit verifies a relationship, not performance. An approved qualitative case verifies the described public observation, not private traffic or revenue. Retention policies should preserve immutable baselines, versioned logic, raw receipts, and publication permissions long enough to audit comparisons. Deletion or access restrictions should follow contractual and legal requirements. The result is an evidence surface that is useful to buyers and answer engines without turning operational data into an uncontrolled public dump.
- Document public, client-only, and service-role-only fields before the first run.
- Keep redactions visible in the methodology and consistent across observations.
- Never remove negative evidence to make a public result look stronger.
- Bind each customer statement and metric to current publication permission.
- Retain immutable baseline artifacts, logic versions, hashes, and approval records.
How audit findings become remediation priorities
The audit should produce a decision ledger, not a generic checklist. Each finding identifies the affected audience and prompt family, the page or external evidence involved, the observed failure, the proposed change, the responsible owner, the release gate, and the measurement date. Technical blocks and contradictory identity records come first because they can invalidate every downstream page. Next come overlapping commercial URLs, missing source ledgers, weak direct answers, and unsupported claims. External profiles, reviews, original research, and earned coverage follow on their real timelines rather than being represented as completed site work. Expected impact is a prioritization hypothesis, not a guaranteed score gain. After a change ships, production proof confirms the intended HTML, schema, links, analytics, and robots state. Only the unchanged prompt panel can show whether answer visibility moved, and only persisted lead and revenue records can show whether the movement produced clients.
- Bind every finding to a page, prompt family, evidence gap, owner, and verification date.
- Repair crawl, identity, and canonical contradictions before increasing publication volume.
- Separate site-controlled work from authenticated profiles and third-party editorial outcomes.
- Verify the live release before rerunning the unchanged measurement panel.
- Use observed prompt and pipeline movement to reprioritize the next cycle.
Results at the level the evidence supports.
Swipe or scroll the evidence table horizontally to inspect every field.
| Evidence | Result | Method | Date | Limitation and source |
|---|---|---|---|---|
| 10xSearch fixed-panel baseline | The first published baseline contains 100 prompt-engine observations, 100 raw hashed captures, 39 responses with at least one cited URL, 0 recommendations, and 0 errors. | The versioned growth-audit-50.v1 panel submitted 50 fixed national prompts independently to Gemini and Claude. Every observation retained the model, United States API execution scope, raw answer, citations, verdicts, timestamp, and SHA-256 capture hash. | Measured August 13, 2026 | This is a two-engine point-in-time API baseline. A cited URL does not mean 10xSearch was mentioned or recommended, and consumer interfaces or other locations can return different answers. Inspect the evidence |
| Doug Leibinger AI visibility | 40 questions produced strict recognition in the newest August 10 batch. The durable win ledger contains 43 won questions, with 3 older wins labeled historical rather than current. | 59 questions were checked across six engines. A question counts as current only when at least one engine names Doug in answer prose in the single newest batch. | Measured August 10, 2026 | The legacy panel did not persist provider-model or execution-location fields. The public asset labels those fields not recorded instead of inferring them. Inspect the evidence |
How the conclusion is produced.
- 1Define the business category, intended audience, deliberate geographic scope, and name aliases from source evidence.
- 2Freeze the prompt-panel version before any provider request runs.
- 3Submit each prompt independently to each configured engine and retain the raw provider response.
- 4Apply deterministic name, negative-signal, position, sentiment, and citation extraction rules.
- 5Persist the run and each prompt-engine observation in service-role-only tables before treating the report as fulfilled.
- 6Render methodology, limitations, prompts, captures, citations, and hashes into the report so the public number can be checked.
- 7Repeat the same panel after material work is live, and compare only compatible panel versions and verdict contracts.
What this page does not prove.
- No agency controls whether an answer engine cites a specific page on a specific date.
- AI answers vary by engine, model, interface, account context, date, and location.
- Recognition, recommendation, citation, ranking position, sentiment, traffic, and revenue are separate measures and should not be blended into one score.
- Case results show what happened for the named client in the stated window. They are not a guarantee of the same outcome for another business.
- API answers may differ from a consumer product interface, especially when the consumer product uses account history, personalization, or a different retrieval layer.
- A deterministic text verdict can be audited, but it is still an operational classification rather than access to the provider's internal ranking logic.
- Comparisons across panel versions require an explicit bridge analysis. The audit does not silently blend incompatible baselines.
Choose by operating complexity.
Public starting prices as reviewed 2026-08-13. Final scope depends on markets, entities, evidence readiness, integrations, and publishing requirements.
For a brand with a workable technical foundation that needs the AI visibility audit, schema audit, 60-day asset plan, publishing velocity, and ongoing monitoring.
For a network leader or established brand that also needs a higher-touch entity graph, authority-source expansion, and press-placement scoping.
For a principal who wants maximum founder involvement in positioning, category framing, and competitive response, with the lifetime monthly rate described on the pricing page.
Direct answers before a sales call.
Why use 50 prompts?
+
A 50-prompt panel is large enough to cover several commercial decision families without letting one prompt dominate the result, while remaining small enough to rerun consistently and inspect at the raw-answer level.
Why not generate prompts with AI for every audit?
+
Generated prompts can improve discovery, but they weaken repeatability. The scored panel is fixed. New candidate prompts can be researched separately and admitted only through a versioned panel change.
Does a citation count as a recommendation?
+
No. A source URL can be cited without the business being recommended. Citation, strict recognition, recommendation, position, and sentiment are tracked separately.
How is location handled?
+
The audit stores prompt market and execution location separately. A city appears in the scored panel only when the business deliberately wants that local market. For 10xSearch, Chicago is not a primary scored market.
Can I get the raw data?
+
Yes. New reports render the raw captures and preserve database receipts. Approved public case assets can also expose machine-readable JSON.
Follow the source, not the adjective.
Inspect the narrower operating questions.
What should you inspect next?
Audit the questions you intend to win.
Use a stable panel, keep the raw receipts, and make every headline traceable to a dated observation.
Or book a fit call