Immutable methodology · ai-discoverability-rubric-v3

FutureContent AI Discoverability

AI Discoverability is a technical assessment of observable machine access, identity, metadata, structured representation, provenance, and citation resolvability for supported AI search and answer products.

Pinned identities

Evaluation Version
page-scan-2026-08-18-v16
Manifest
assessment-e79ae98ad2b433fc
Rubric
ai-discoverability-rubric-v3
Specification
page-scan-v1-specification-v1
Ownership Registry
page-scan-v1-ownership-v1
Provider Registry
page-scan-v1-provider-registry-v2
Weights
page-scan-157-check-scoring-v1

The stable AI Discoverability methodology page points here while this is the active version.

Technical ownership boundary

AI Discoverability asks whether supported machines can retrieve the page, identify its canonical and language context, interpret its directives and structured representation, and resolve attached citations. It does not judge search intent, answer quality, readability, factual truth, or whether a provider will include, cite, recommend, or rank the page.

Main content hidden behind login, OAuth, consent, CAPTCHA, interaction, or unresolved client rendering creates one access-gate finding and can make dependent semantic checks unavailable. Multi-step agent operability remains outside this rubric.

Supported scored discovery products

ChatGPT

OAI-SearchBot · 3.6 points within the capped access component

Robots permission is observable eligibility, not proof of inclusion or access through verified OpenAI infrastructure.

OpenAI crawler controls · verified 2026-08-13

Google Gemini

Googlebot · 3.6 points within the capped access component

This mapping covers AI features in Google Search, whose access control is Googlebot; it is not a product-specific Gemini crawler.

Google AI features and websites · verified 2026-08-13

Claude

Claude-SearchBot · 3.6 points within the capped access component

Permission supports search visibility but cannot prove actual retrieval or citation.

Anthropic crawler controls · verified 2026-08-13

Microsoft Copilot

bingbot · 3.6 points within the capped access component

Use bingbot access with applicable Microsoft-supported snippet and presentation controls; eligibility does not guarantee a Copilot citation.

Microsoft controls for Bing AI experiences · verified 2026-08-13

Perplexity AI

PerplexityBot · 3.6 points within the capped access component

Permission alone cannot prove WAF passage from verified Perplexity infrastructure or inclusion in results.

Perplexity crawlers · verified 2026-08-13

Implicit permission and an explicit Allow are equivalent. An unavailable product preserves its share in a score range; the other products are not reweighted. Shared restrictions cannot deduct more than the 18-point access cap.

Training controls and unsupported products

Visible non-scoring controls

  • ChatGPT: ChatGPT-User controls user triggered retrieval. User-triggered fetches are not automatic discovery and robots.txt may not apply.
  • ChatGPT: GPTBot controls training. Training preference is independent from ChatGPT Search discovery.
  • Google Gemini: Google-Extended controls training. Google-Extended is a control token without a separate HTTP user agent and does not affect Google Search inclusion or ranking.
  • Claude: Claude-User controls user triggered retrieval. User-directed retrieval is distinct from search indexing.
  • Claude: ClaudeBot controls training. Training preference is independent from Claude search discovery.
  • Perplexity AI: Perplexity-User controls user triggered retrieval. Perplexity documents that this user-requested fetcher generally ignores robots.txt; it cannot represent discovery eligibility.

Blocking a training crawler does not reduce AI Discoverability. A FutureContent request with a provider user-agent is a non-scoring delivery simulation and cannot prove access from verified provider infrastructure.

Unsupported or unverifiable products

  • Grok: visible, non-scoring, and not mapped until an authoritative first-party product-specific crawler contract can be verified.
  • Meta AI: visible, non-scoring, and not mapped until an authoritative first-party product-specific crawler contract can be verified.
  • DeepSeek: visible, non-scoring, and not mapped until an authoritative first-party product-specific crawler contract can be verified.
  • Mistral Vibe (formerly Le Chat): visible, non-scoring, and not mapped until an authoritative first-party product-specific crawler contract can be verified.

llms.txt orientation evidence

The canonical root /llms.txt is one proprietary FutureContent agent-orientation signal worth at most two points. It must be retrievable Markdown, identify the site, and link to relevant same-site authoritative resources that resolve. Missing, invalid, unusable, materially stale, or misleading outcomes share one maximum deduction and do not stack.

The bounded outcome is cached by origin for 24 hours with its fetch timestamp and resolved URL; a user-requested recheck bypasses reuse. This is not crawl permission, training consent, inclusion, citation, recommendation, or ranking evidence. It does not affect Google Search visibility or ranking; Google Search gives llms.txt no special treatment.

Owned scored checks

These current Defect Families are derived from the pinned Ownership Registry. Presence, syntax, access, and delivery belong here; SEO and AI Content own the meaning and usefulness of successfully acquired content.

AI Discoverability defect families (16)

  1. Supported discovery-product access

    robots.txt acquisition and effective policy for each scored discovery token.

  2. Valid robots.txt syntax

    Mobile-preferred Lighthouse robots-txt audit result.

  3. Indexing and link-following eligibility

    HTTP X-Robots-Tag, initial HTML, rendered metadata, and link-level directives.

  4. Search and answer snippet eligibility

    Robots metadata, X-Robots-Tag, and bounded rendered data-nosnippet coverage.

  5. Main-content machine access

    Normalized rendered main content and detected login, consent, CAPTCHA, interaction, or client-rendering gates.

  6. Required self-referencing canonical

    Submitted URL, redirects, final normalized URL, HTTP Link headers, initial head, and rendered head.

  7. Applicable language and region alternates

    Credible alternate pages, rendered hreflang declarations, targets, and reciprocal cluster evidence.

  8. Title presence and parseability

    Initial and rendered title elements.

  9. Meta-description presence and parseability

    Initial and rendered meta-description elements.

  10. Compatible Schema.org primary entity

    Rendered JSON-LD, Microdata, and RDFa primary-entity candidates.

  11. Structured-data syntax and supported vocabulary

    Every bounded rendered structured-data block and supported Schema.org vocabulary snapshot.

  12. Open Graph completeness

    Rendered og:title, og:description, og:url, and bounded og:image response evidence.

  13. Applicable provenance and freshness signals

    Visible and structured authorship, publisher, dates, version, owner, and responsible organization.

  14. Citation attachment and technical resolvability

    Existing citation attachment plus safe bounded target status, redirects, and content type.

  15. Crawlable and resolvable link targets

    Rendered link destinations and safe bounded resolution evidence.

  16. Useful root llms.txt

    Canonical root /llms.txt status, resolved URL, Markdown structure, site identity, and bounded same-site resource resolution.

AI Discoverability contributes 25% to a complete Page Optimization Score. Unavailable checks retain their possible deductions in a non-reweighted score range.

Limitations and no guarantees

This single-run technical assessment does not guarantee provider crawling, inclusion, citation, recommendation, ranking, training behavior, trustworthiness, or end-to-end agent completion.