AI Search Glossary
Plain-language definitions of GEO, AEO, AI crawlers, structured data, and measurement terms used in AI visibility work. Built as a neutral reference—and as a durable internal-linking layer for tools and articles.
173 terms · free reference · no signup
Showing 173 of 173
A
- AEO (Answer Engine Optimization)AEO is optimizing content so answer engines—systems that return direct answers rather than only links—can find, understand, and surface your information.Core concepts
- Affiliate DisclosureAn affiliate disclosure states that a publisher may earn commissions from links or recommendations, informing readers of that commercial relationship.Core concepts
- AI CrawlerAn AI crawler is a bot that fetches public web pages for an AI company or dataset—often for training, indexing, or live assistant browsing.Crawlers & bots
- AI OverviewAn AI Overview is Google’s generated summary shown above traditional results for some queries, often with links to supporting sources.Core concepts
- AI Referral TrafficAI referral traffic is website visits whose referrer is an AI product such as ChatGPT, Perplexity, or Gemini.Measurement
- AI SearchAI search is the use of AI systems—chatbots, AI Overviews, and answer engines—to find information, products, or brands instead of only classic link lists.Core concepts
- AI Search TrafficAI search traffic is visit volume associated with AI-driven discovery—referrals from answer engines or clicks from AI-enhanced search features.Measurement
- AI Training Opt-OutAI training opt-out is the practice of signaling or enforcing that public site content should not be used to train machine learning models.Crawlers & bots
- AI VisibilityAI visibility is how often and how prominently a brand appears in AI-generated answers, citations, and recommendations across major AI products.Core concepts
- Allow Directive (robots.txt)Allow is a robots.txt rule that permits a crawler to fetch a path, often used to carve exceptions inside a broader Disallow.Crawlers & bots
- Alternative PageAn alternative page targets buyers considering substitutes for a named product—often ranked as “best alternatives to X.”Core concepts
- Anchor TextAnchor text is the clickable wording of a hyperlink, signaling to users and machines what the destination page is about.Technical
- Answer EngineAn answer engine is a system that returns a direct answer or synthesis to a query, often with citations, instead of only a ranked list of links.Core concepts
- Answer Engine ResultsAnswer engine results are the synthesized answers, citations, and follow-ups an answer engine returns instead of only a ranked list of blue links.Core concepts
- Answer ShareAnswer share is the portion of generative answers in a sample that include a brand as a recommendation, mention, or cited source.Measurement
- ApplebotApplebot is Apple’s web crawler used for features such as Safari and Spotlight search, with related tokens for generative AI preferences.Crawlers & bots
- Assisted ConversionAn assisted conversion is a conversion where an earlier AI or marketing touch influenced the path even if it was not the last click.Measurement
B
- Bing CopilotBing Copilot is the generative assistant experience in Microsoft Bing and Edge that answers queries with AI summaries, links, and follow-up chat.Core concepts
- Bot ManagementBot management is the set of controls—robots.txt, WAF rules, rate limits, challenges—used to identify and shape automated traffic.Crawlers & bots
- Brand Authority ScoreA brand authority score is a numeric estimate of how strong and trusted a brand’s entity appears across the web and search systems.Measurement
- Brand SERPsBrand SERPs are the search results that appear for a brand-name query, including the official site, profiles, news, and knowledge features.Core concepts
- Branded Search LiftBranded search lift is an increase in search volume or clicks for a brand’s name queries after exposure on other channels, including AI answers.Measurement
- Breadcrumb SchemaBreadcrumb schema marks the hierarchical path of a page—Home › Category › Item—so search engines can show and understand site structure.Technical
- Byline StrategyByline strategy is the deliberate choice of named authors, bios, and expert placement to build credibility and entity signals for content.Core concepts
- BytespiderBytespider is a ByteDance web crawler observed collecting public pages, often associated with AI and content products.Crawlers & bots
C
- Canonical URLA canonical URL is the preferred address for a piece of content when duplicates or near-duplicates exist, signaled mainly via rel=canonical.Technical
- CCBotCCBot is Common Crawl’s crawler that builds open web datasets widely reused by researchers and AI developers.Crawlers & bots
- CDN CachingCDN caching stores copies of responses at edge locations so users and bots receive content from nearby servers instead of only the origin.Technical
- Changelog SEOChangelog SEO is optimizing public release notes so product changes are discoverable, crawlable, and citable by search and AI systems.Core concepts
- ChatGPTChatGPT is OpenAI’s conversational AI product that answers questions, writes text, and uses tools or browsing depending on the plan and mode.AI models
- Citation (AI answers)In AI answers, a citation is a reference or link to a source the system used or attributed when generating a response.Core concepts
- Citation CardA citation card is a UI element in AI answers that shows a source title, link, or snippet attributing part of the generated response.Core concepts
- Citation RateCitation rate is the share of AI answers or prompt runs in which a brand’s URL or named source is explicitly attributed.Measurement
- ClaudeClaude is Anthropic’s family of large language models and chat products used for writing, analysis, coding, and research-style assistance.AI models
- ClaudeBotClaudeBot is Anthropic’s crawler for collecting public web content that may support Claude model training and related improvements.Crawlers & bots
- CLS (Cumulative Layout Shift)Cumulative Layout Shift measures how much visible content unexpectedly moves during the page’s life as elements load or change size.Technical
- Community SourceA community source is user-generated discussion—forums, Discord, Reddit, GitHub issues—that search and AI systems may treat as evidence about a product.Core concepts
- Comparison PageA comparison page evaluates two or more products, approaches, or plans side by side so buyers can choose among realistic alternatives.Core concepts
- Competitive DisplacementCompetitive displacement is when an AI answer recommends or cites a rival in a slot where your brand previously appeared or should compete.Measurement
- Content ClusterA content cluster is a group of related pages organized around a pillar topic, interlinked so users and crawlers grasp the subject as a whole.Core concepts
- Content DecayContent decay is the gradual loss of traffic, rankings, or AI visibility for a page as it becomes outdated or outcompeted.Measurement
- Content FreshnessContent freshness is how recently a page’s substantive facts were updated, and how clearly that recency is visible to people and machines.Measurement
- Content Licensing (AI)Content licensing for AI is the legal and contractual framework governing how publishers’ text and media may be used to train or ground models.Crawlers & bots
- Content Security Policy (CSP)Content Security Policy is an HTTP security header that restricts which scripts, styles, images, and other resources a page may load.Technical
- Context WindowA context window is the maximum span of tokens—prompt, history, and retrieved text—a language model can consider in one generation pass.Technical
- Copilot (Microsoft)Copilot is Microsoft’s brand for AI assistants embedded in Bing, Windows, Edge, Microsoft 365, and developer tools that help users complete tasks.AI models
- Core Web VitalsCore Web Vitals are Google’s primary field metrics for loading, interactivity, and visual stability: LCP, INP, and CLS.Technical
- Crawl BudgetCrawl budget is the practical limit on how many URLs a search engine will crawl on a site in a given period, shaped by demand and capacity.Technical
- Crawl-delayCrawl-delay is an unofficial robots.txt directive that asks a crawler to wait a stated number of seconds between successive requests.Crawlers & bots
- Custom GPTA Custom GPT is a tailored ChatGPT assistant configured with instructions, knowledge files, and optional actions for a specific use case.AI models
D
- Data StudyA data study is a published analysis of a dataset that answers a defined question with methods, findings, and usually visual summaries.Core concepts
- Developer Relations (DevRel)Developer relations is the practice of educating, supporting, and advocating for developers who use or evaluate a technical product.Core concepts
- Digital PRDigital PR earns online coverage, links, and mentions through newsworthy stories, data, and relationships rather than only buying placements.Core concepts
- Digraph of ContentA digraph of content is the directed graph formed by pages as nodes and hyperlinks as one-way edges between them on a site or corpus.Technical
- Directory CompletenessDirectory completeness measures how fully and accurately a brand’s listings are filled on relevant directories and profiles.Measurement
- Disallow Directive (robots.txt)Disallow is a robots.txt rule that asks matching crawlers not to fetch URLs under a given path prefix.Crawlers & bots
- Docs as MarketingDocs as marketing treats product documentation as a growth surface—earning traffic, trust, and AI citations—not only a post-sale help resource.Core concepts
- Doorway PageA doorway page is low-value content built mainly to rank for similar queries and funnel users to a single destination, often in many near-duplicate variants.Core concepts
E
- Edge WorkerAn edge worker is lightweight code that runs on CDN or edge nodes to modify requests and responses near the user before or after cache.Technical
- EEAT (Experience, Expertise, Authoritativeness, Trust)EEAT is Google’s quality framework emphasizing experience, expertise, authoritativeness, and trustworthiness in content and its creators.Core concepts
- Entity ConsistencyEntity consistency means presenting the same core facts about a brand—name, description, products, and relationships—across public sources machines compare.Core concepts
- Entity DriftEntity drift is divergence over time in how public sources describe a brand’s identity, products, or relationships.Measurement
- Entity HomeAn entity home is the primary public URL that should represent an entity—usually the official About or organization page machines treat as canonical identity.Core concepts
- Entity SEOEntity SEO is optimizing how search engines and AI systems understand your brand as a distinct entity—who you are, what you offer, and how you relate to other entities.Core concepts
- Experience SignalAn experience signal is evidence that content reflects first-hand use, observation, or practice rather than only second-hand summary.Core concepts
- Expert QuoteAn expert quote is an attributed statement from a qualified person used in journalism, research, or branded content to add authority.Core concepts
F
- Faceted NavigationFaceted navigation lets users filter listings by attributes such as size, color, or price, often creating many URL combinations.Technical
- Fact ConsistencyFact consistency is agreement on core claims—pricing tiers, features, founding details—across owned pages and major third-party sources.Measurement
- FAQPage SchemaFAQPage schema marks a page of question-and-answer pairs so machines can parse each question and its accepted answer as structured data.Technical
- FirecrawlFirecrawl is a web crawling and extraction product that turns pages into clean, LLM-ready data such as Markdown or structured JSON.Crawlers & bots
G
- Gem (Gemini)A Gem is a customized Gemini assistant with saved instructions and optional knowledge so users can reuse a specialized chat persona.AI models
- GeminiGemini is Google’s family of multimodal AI models and related consumer and workspace products that generate answers, media, and assistant responses.AI models
- Generative UIGenerative UI is an interface that AI systems assemble or adapt dynamically—components, layouts, or widgets—based on the user’s query and context.Technical
- GEO (Generative Engine Optimization)GEO is the practice of improving how often and how accurately AI systems mention, cite, or recommend your brand in generative answers.Core concepts
- GEO ScoreA GEO score is a composite estimate of how ready and visible a brand is for generative and answer-engine discovery.Measurement
- Google AI ModeGoogle AI Mode is a Search experience that answers queries with conversational, multi-step AI responses inside Google rather than only classic result lists.Core concepts
- Google-ExtendedGoogle-Extended is a robots.txt product token used to express preferences about using site content for Gemini and certain Google AI features.Crawlers & bots
- GooglebotGooglebot is Google’s primary family of crawlers that fetch pages for Google Search discovery, indexing, and related search features.Crawlers & bots
- GPTBotGPTBot is OpenAI’s web crawler user-agent used for collecting public content that may support OpenAI models and products.Crawlers & bots
- GroundingGrounding is connecting an AI model’s output to retrieved external information—documents, search results, or tools—so answers stay tied to sources.Core concepts
H
- Hallucination (AI)In AI systems, a hallucination is a confident-sounding output that is fabricated or not supported by reliable sources.Core concepts
- Helpful ContentHelpful content is material created primarily for people that satisfies intent with original, reliable substance rather than search-engine-first filler.Core concepts
- HowTo SchemaHowTo schema describes a multi-step instructional process—tools, supplies, and ordered steps—so machines can understand task guidance.Technical
- hreflanghreflang is HTML or HTTP markup that tells search engines which language and regional URL variant is intended for which audience.Technical
- HSTS (HTTP Strict Transport Security)HSTS is a response header that tells browsers to use HTTPS only for a host, blocking insecure HTTP access after the first secure visit.Technical
- HydrationHydration is the process of attaching client-side JavaScript to server-rendered HTML so a static page becomes an interactive application.Technical
I
- IndexabilityIndexability is whether a URL is allowed and suitable to be stored in a search engine’s index and potentially shown in results.Technical
- INP (Interaction to Next Paint)Interaction to Next Paint measures how quickly a page visually responds after user interactions such as taps, clicks, and key presses.Technical
- Integration PageAn integration page documents how a product connects with another system—setup steps, scopes, limits, and supported workflows.Core concepts
- Internal LinkingInternal linking is the practice of connecting pages on the same site with hyperlinks to guide users, distribute equity, and clarify structure.Technical
- IP AllowlistAn IP allowlist is a set of trusted address ranges permitted through a firewall or WAF without the same blocks applied to unknown clients.Crawlers & bots
J
K
- Knowledge GraphA knowledge graph is a network of entities and relationships that machines use to understand how people, brands, and concepts connect.Core concepts
- Knowledge PanelA knowledge panel is a search results module that summarizes an entity—brand, person, or place—using knowledge-graph style facts and links.Core concepts
L
- Large Language Model (LLM)A large language model is a neural network trained on vast text data to predict and generate human-like language.AI models
- lastmodlastmod is an optional sitemap field that states when a listed URL’s content last changed, helping crawlers prioritize recrawls.Technical
- LCP (Largest Contentful Paint)Largest Contentful Paint measures how long it takes for the largest visible content element in the viewport to finish rendering.Technical
- Live Browsing AgentA live browsing agent is a user-agent that fetches pages on demand during an assistant conversation rather than bulk-training the open web.Crawlers & bots
- llms.txtllms.txt is a proposed root-level markdown file that gives AI systems a curated summary of a site’s important pages and context.Technical
- Local Pack AILocal Pack AI refers to AI-enhanced local results—maps packs, summaries, and assistant answers that recommend nearby businesses for intent-rich queries.Core concepts
- LocalBusiness SchemaLocalBusiness schema is Schema.org markup for a physical or service-area business, including address, geo, hours, and contact details.Technical
- Log File AnalysisLog file analysis examines server or CDN request logs to see which crawlers hit which URLs, status codes, and response patterns over time.Technical
M
- MarkdownMarkdown is a lightweight plain-text format that uses simple symbols for headings, lists, and links, often rendered as HTML.Technical
- Markdown for LLMsMarkdown for LLMs is the practice of offering clean, structured Markdown so language models and crawlers can parse headings, lists, and links easily.Technical
- Media KitA media kit is a package of approved brand assets, facts, and contacts that help journalists and partners cover an organization accurately.Core concepts
- Mention SentimentMention sentiment is the positive, neutral, or negative framing of a brand when it appears inside an AI-generated answer.Measurement
- Meta RobotsMeta robots is an HTML meta tag that gives crawlers page-level instructions such as index/noindex and follow/nofollow.Technical
- MicrodataMicrodata is an HTML syntax that annotates visible elements with Schema.org types and properties using attributes like itemscope and itemprop.Technical
- Migration GuideA migration guide explains how to move data, workflows, or users from one tool or version to another with sequenced steps and risks.Core concepts
- Model DriftModel drift is change in an AI system’s outputs over time for the same prompts, caused by updates, tooling, or data shifts.Measurement
- Multimodal SearchMultimodal search lets users query with more than text—images, voice, video, or mixed inputs—and receive results that understand those modalities.Core concepts
N
- NAP ConsistencyNAP consistency means keeping name, address, and phone identical across a business’s website, profiles, and local citations.Core concepts
- nofollownofollow is a link or robots hint that asks crawlers not to treat a link as an editorial endorsement for ranking credit.Technical
- noindexnoindex is a robots directive that asks compliant crawlers not to include a URL in their search index.Technical
O
- OAI-SearchBotOAI-SearchBot is OpenAI’s crawler for indexing public content used in ChatGPT search and retrieval-style features.Crawlers & bots
- Organization SchemaOrganization schema is Schema.org markup that describes a company or brand—name, URL, logo, contact points, and sameAs profile links.Technical
- Original ResearchOriginal research is primary data or analysis a publisher produces—surveys, experiments, or novel datasets—not merely a rewrite of others’ findings.Core concepts
- Orphan PageAn orphan page is a URL with no internal links from other pages on the same site, making it hard for users and crawlers to discover.Technical
P
- PaginationPagination splits a long list of items across multiple URLs—page 2, page 3, and so on—with navigation between them.Technical
- PerplexityPerplexity is an AI-powered answer engine that synthesizes web results into cited responses rather than only listing ranked links.AI models
- PerplexityBotPerplexityBot is Perplexity’s web crawler for discovering and indexing public pages used in Perplexity search and answers.Crawlers & bots
- Pillar PageA pillar page is a comprehensive hub article that covers a core topic broadly and links to deeper cluster content on subtopics.Core concepts
- Press BoilerplateA press boilerplate is a short standard paragraph describing an organization for the end of releases and media kits.Core concepts
- Primary SourceA primary source is original evidence—documents, data, official statements, or firsthand records—rather than someone else’s interpretation.Core concepts
- Product SchemaProduct schema is Schema.org markup that describes a sellable item—name, offers, identifiers, availability, and optional reviews.Technical
- Product-Led ContentProduct-led content teaches or solves problems using the real product experience—workflows, screens, and outcomes—rather than only brand messaging.Core concepts
- Programmatic SEOProgrammatic SEO generates many pages from templates and datasets to cover structured query patterns at scale, with quality controls.Core concepts
- Project KnowledgeProject knowledge is a scoped set of documents or notes attached to an AI workspace so answers prefer that corpus for a task or team.Technical
- PromptA prompt is the input text or instruction a user or system gives a language model to steer its response.Measurement
- Prompt BatteryA prompt battery is a fixed, versioned set of test questions used repeatedly to measure brand presence in AI answers over time.Measurement
Q
R
- Readiness ChecklistA readiness checklist is a structured audit of technical, content, and entity items that support AI search and classic SEO discoverability.Measurement
- Render-Blocking ResourcesRender-blocking resources are CSS, scripts, or fonts the browser must process before it can paint meaningful page content for the user.Technical
- Retrieval-Augmented Generation (RAG)RAG is an architecture that retrieves relevant documents at query time and conditions a language model’s answer on that retrieved context.Technical
- Review VelocityReview velocity is the rate at which new customer reviews appear on major platforms over a given period.Measurement
- robots.txtrobots.txt is a text file at a site’s root that suggests which crawlers may fetch which paths, using User-agent and Allow/Disallow rules.Technical
S
- sameAssameAs is a Schema.org property listing official URLs that represent the same entity, such as social profiles or encyclopedia entries.Technical
- Scaled ContentScaled content is publishing at high volume—often with automation or large writer pools—to cover many keywords or entities quickly.Core concepts
- Schema.orgSchema.org is a collaborative vocabulary of types and properties for describing things on the web in a machine-readable way.Technical
- Secondary SourceA secondary source interprets, summarizes, or synthesizes primary evidence—reviews, explainers, textbooks, and most news analysis.Core concepts
- Server-Side Rendering (SSR)Server-side rendering generates full HTML on the server for each request so browsers and crawlers receive content without waiting on client JavaScript.Technical
- Share of ModelShare of model is how often a brand appears in answers from a specific AI system relative to competitors on a fixed prompt set.Measurement
- Share of Voice (AI answers)In AI answers, share of voice is how often your brand appears among recommendations or mentions relative to competitors for a set of prompts.Measurement
- Shopping GraphA shopping graph is a structured network of products, merchants, prices, and attributes that power product discovery in search and AI shopping features.Core concepts
- SitemapA sitemap is a file—usually XML—that lists important URLs on a site to help crawlers discover pages efficiently.Technical
- Sitemap Directive (robots.txt)The Sitemap directive in robots.txt declares the absolute URL of an XML sitemap so crawlers can discover the site’s URL list.Crawlers & bots
- Soft 404A soft 404 is a URL that returns HTTP 200 OK but shows “not found” or empty content, misleading crawlers that rely on status codes.Technical
- Source DiversitySource diversity is the variety of independent publishers and page types that discuss or corroborate a brand or claim.Measurement
- Sponsored ContentSponsored content is material paid for or materially influenced by an advertiser, presented within a publisher’s editorial-looking environment.Core concepts
- Structured DataStructured data is machine-readable markup—often JSON-LD—that describes a page’s content using a shared vocabulary such as Schema.org.Technical
- Survey MethodologySurvey methodology is the design of how respondents are sampled, questioned, and analyzed so results can be interpreted responsibly.Core concepts
- System PromptA system prompt is the hidden or privileged instruction layer that steers a model’s role, rules, and style before the user’s messages are applied.AI models
T
- Temperature (LLM sampling)Temperature is a generation setting that controls how randomly a language model samples next tokens—lower is more deterministic, higher is more varied.AI models
- Thin ContentThin content is page substance too shallow, unoriginal, or unhelpful to satisfy users—or to deserve strong search and AI visibility.Core concepts
- Third-Party CorroborationThird-party corroboration is independent confirmation of a brand’s claims on sites the brand does not control.Measurement
- Thought LeadershipThought leadership is expert commentary that advances how an industry understands a problem, not only promotes a vendor’s features.Core concepts
- Tool Use (AI agents)Tool use is when a language model calls external functions—search, code execution, APIs, or browsers—to gather data or take actions while answering.Technical
- Topical AuthorityTopical authority is the depth and credibility a site demonstrates on a subject through comprehensive, consistent, well-linked coverage.Core concepts
- Training CrawlerA training crawler is a bot that collects public web content primarily to build or improve machine learning training datasets and models.Crawlers & bots
- Training DataTraining data is the text, code, and other media used to teach a machine learning model its parameters before deployment.AI models
U
- UGC SEOUGC SEO is optimizing and moderating user-generated content—reviews, forums, Q&A—so it helps discovery without creating spam or trust problems.Core concepts
- Uncontested PromptAn uncontested prompt is a test or real-world query where no relevant competitor is mentioned—only one brand, or none, appears.Measurement
- User-AgentA User-Agent is an identifier string a client sends to a server to describe the browser, bot, or application making the request.Crawlers & bots
V
- Versus Page (A vs B)A versus page is a focused comparison of exactly two named options, usually titled in an “A vs B” pattern for decision-stage queries.Core concepts
- View-Through BrandView-through brand impact is brand demand created when users see a brand in an AI answer or ad without clicking it immediately.Measurement
- Visibility ProbeA visibility probe is a single controlled test—usually one prompt on one model—to sample whether and how a brand appears in the answer.Measurement
- Voice Assistant SearchVoice assistant search is information retrieval through spoken queries to assistants such as smart speakers, phones, or in-car systems that read answers aloud.Core concepts
W
- WAF Bot ScoreA WAF bot score is a numeric or categorical signal from a web application firewall estimating how likely a request is automated versus human.Crawlers & bots
- WikidataWikidata is a free, collaborative knowledge base of structured items and statements that machines and Wikimedia projects reuse as linked data.Core concepts
- Wikipedia SEOWikipedia SEO is the practice of earning and maintaining accurate, well-sourced encyclopedia coverage that strengthens public entity understanding.Core concepts
X
Z
- Zero-Click RateZero-click rate is the share of searches or AI sessions where users get an answer without clicking through to a publisher’s website.Measurement
- Zero-Click SearchA zero-click search is a query where the user gets the answer on the results page itself and does not click through to a website.Core concepts
All terms remain in the page HTML for crawlers; search and filters only hide rows in the browser.
Suggest a term → Missing something? Email a term and a one-line definition draft.
Frequently asked questions
- SEO traditionally optimizes for ranked links in search engines. GEO (Generative Engine Optimization) focuses on how AI systems mention, cite, or recommend your brand in generated answers. They share foundations—crawlable content, authority, clear entities—but GEO adds answer-engine and citation-oriented tactics.