Immutable methodology · ai-discoverability-rubric-v3
FutureContent AI Discoverability
AI Discoverability is a technical assessment of observable machine access, identity, metadata, structured representation, provenance, and citation resolvability for supported AI search and answer products.
Pinned identities
- Evaluation Version
page-scan-2026-08-18-v16- Manifest
assessment-e79ae98ad2b433fc- Rubric
ai-discoverability-rubric-v3- Specification
page-scan-v1-specification-v1- Ownership Registry
page-scan-v1-ownership-v1- Provider Registry
page-scan-v1-provider-registry-v2- Weights
page-scan-157-check-scoring-v1
The stable AI Discoverability methodology page points here while this is the active version.
Technical ownership boundary
AI Discoverability asks whether supported machines can retrieve the page, identify its canonical and language context, interpret its directives and structured representation, and resolve attached citations. It does not judge search intent, answer quality, readability, factual truth, or whether a provider will include, cite, recommend, or rank the page.
Main content hidden behind login, OAuth, consent, CAPTCHA, interaction, or unresolved client rendering creates one access-gate finding and can make dependent semantic checks unavailable. Multi-step agent operability remains outside this rubric.
Supported scored discovery products
ChatGPT
OAI-SearchBot · 3.6 points within the capped access component
Robots permission is observable eligibility, not proof of inclusion or access through verified OpenAI infrastructure.
OpenAI crawler controls · verified 2026-08-13
Google Gemini
Googlebot · 3.6 points within the capped access component
This mapping covers AI features in Google Search, whose access control is Googlebot; it is not a product-specific Gemini crawler.
Google AI features and websites · verified 2026-08-13
Claude
Claude-SearchBot · 3.6 points within the capped access component
Permission supports search visibility but cannot prove actual retrieval or citation.
Anthropic crawler controls · verified 2026-08-13
Microsoft Copilot
bingbot · 3.6 points within the capped access component
Use bingbot access with applicable Microsoft-supported snippet and presentation controls; eligibility does not guarantee a Copilot citation.
Microsoft controls for Bing AI experiences · verified 2026-08-13
Perplexity AI
PerplexityBot · 3.6 points within the capped access component
Permission alone cannot prove WAF passage from verified Perplexity infrastructure or inclusion in results.
Perplexity crawlers · verified 2026-08-13
Implicit permission and an explicit Allow are equivalent. An unavailable product preserves its share in a score range; the other products are not reweighted. Shared restrictions cannot deduct more than the 18-point access cap.
Training controls and unsupported products
Visible non-scoring controls
- ChatGPT:
ChatGPT-Usercontrols user triggered retrieval. User-triggered fetches are not automatic discovery and robots.txt may not apply. - ChatGPT:
GPTBotcontrols training. Training preference is independent from ChatGPT Search discovery. - Google Gemini:
Google-Extendedcontrols training. Google-Extended is a control token without a separate HTTP user agent and does not affect Google Search inclusion or ranking. - Claude:
Claude-Usercontrols user triggered retrieval. User-directed retrieval is distinct from search indexing. - Claude:
ClaudeBotcontrols training. Training preference is independent from Claude search discovery. - Perplexity AI:
Perplexity-Usercontrols user triggered retrieval. Perplexity documents that this user-requested fetcher generally ignores robots.txt; it cannot represent discovery eligibility.
Blocking a training crawler does not reduce AI Discoverability. A FutureContent request with a provider user-agent is a non-scoring delivery simulation and cannot prove access from verified provider infrastructure.
Unsupported or unverifiable products
- Grok: visible, non-scoring, and not mapped until an authoritative first-party product-specific crawler contract can be verified.
- Meta AI: visible, non-scoring, and not mapped until an authoritative first-party product-specific crawler contract can be verified.
- DeepSeek: visible, non-scoring, and not mapped until an authoritative first-party product-specific crawler contract can be verified.
- Mistral Vibe (formerly Le Chat): visible, non-scoring, and not mapped until an authoritative first-party product-specific crawler contract can be verified.
llms.txt orientation evidence
The canonical root /llms.txt is one proprietary FutureContent agent-orientation signal worth at most two points. It must be retrievable Markdown, identify the site, and link to relevant same-site authoritative resources that resolve. Missing, invalid, unusable, materially stale, or misleading outcomes share one maximum deduction and do not stack.
The bounded outcome is cached by origin for 24 hours with its fetch timestamp and resolved URL; a user-requested recheck bypasses reuse. This is not crawl permission, training consent, inclusion, citation, recommendation, or ranking evidence. It does not affect Google Search visibility or ranking; Google Search gives llms.txt no special treatment.
Owned scored checks
These current Defect Families are derived from the pinned Ownership Registry. Presence, syntax, access, and delivery belong here; SEO and AI Content own the meaning and usefulness of successfully acquired content.
AI Discoverability defect families (16)
Supported discovery-product access
robots.txt acquisition and effective policy for each scored discovery token.
Valid robots.txt syntax
Mobile-preferred Lighthouse robots-txt audit result.
Indexing and link-following eligibility
HTTP X-Robots-Tag, initial HTML, rendered metadata, and link-level directives.
Search and answer snippet eligibility
Robots metadata, X-Robots-Tag, and bounded rendered data-nosnippet coverage.
Main-content machine access
Normalized rendered main content and detected login, consent, CAPTCHA, interaction, or client-rendering gates.
Required self-referencing canonical
Submitted URL, redirects, final normalized URL, HTTP Link headers, initial head, and rendered head.
Applicable language and region alternates
Credible alternate pages, rendered hreflang declarations, targets, and reciprocal cluster evidence.
Title presence and parseability
Initial and rendered title elements.
Meta-description presence and parseability
Initial and rendered meta-description elements.
Compatible Schema.org primary entity
Rendered JSON-LD, Microdata, and RDFa primary-entity candidates.
Structured-data syntax and supported vocabulary
Every bounded rendered structured-data block and supported Schema.org vocabulary snapshot.
Open Graph completeness
Rendered og:title, og:description, og:url, and bounded og:image response evidence.
Applicable provenance and freshness signals
Visible and structured authorship, publisher, dates, version, owner, and responsible organization.
Citation attachment and technical resolvability
Existing citation attachment plus safe bounded target status, redirects, and content type.
Crawlable and resolvable link targets
Rendered link destinations and safe bounded resolution evidence.
Useful root llms.txt
Canonical root /llms.txt status, resolved URL, Markdown structure, site identity, and bounded same-site resource resolution.
AI Discoverability contributes 25% to a complete Page Optimization Score. Unavailable checks retain their possible deductions in a non-reweighted score range.
Limitations and no guarantees
This single-run technical assessment does not guarantee provider crawling, inclusion, citation, recommendation, ranking, training behavior, trustworthiness, or end-to-end agent completion.