SEO & Indexing Glossary

Precise definitions of the concepts that decide whether your pages get crawled, indexed, and ranked - written for SEOs, linked to the guides and free tools that act on them.

Indexing Concepts

The systems and protocols that decide whether a URL enters a search engine's index.

Search Engine Indexing

Search engine indexing is the process of analyzing a crawled page's content and storing it in the search engine's index — the database queried when users search. A page that is not indexed cannot rank for any query.

Search Index

A search index is the structured database where a search engine stores the content it has crawled and analyzed. When a user searches, results come from the index — not the live web — which is why unindexed pages are invisible.

Google Indexing API

The Google Indexing API is an official Google API that notifies Google when a URL is added, updated, or deleted, prompting a fast crawl. Google restricts eligible content to job postings and broadcast-event livestream pages.

IndexNow

IndexNow is an open URL-submission protocol, launched by Microsoft Bing and Yandex in 2021, that lets a site notify participating search engines the instant a URL is added, updated, or deleted. Google does not participate.

Structured Data

Structured data is machine-readable markup — written in the schema.org vocabulary, usually as JSON-LD — that labels a page's entities and their relationships so search engines can extract them reliably and consider the page for rich results.

Mobile-First Indexing

Mobile-first indexing is Google's policy of indexing and ranking every site from its mobile version, crawled by Googlebot Smartphone. Content, links, structured data, and robots directives missing from the mobile page are missing from the index.

Index Bloat

Index bloat is the condition where large numbers of a site's low-value URLs — parameter variants, tag archives, thin pages — sit in Google's index, diluting site-wide quality signals and wasting crawl budget on pages that never earn search traffic.

Crawling & Rendering

How search engine bots discover, fetch, and render pages before indexing them.

Crawling

Crawling is the process by which search engine bots discover URLs and fetch their content for analysis. It is the first stage of search visibility: a page must be crawled before it can be rendered, indexed, or ranked.

Googlebot

Googlebot is the web crawler Google uses to discover and fetch pages for its search index. Googlebot Smartphone, which crawls with a mobile user agent, is the primary crawler for nearly all sites under mobile-first indexing.

Crawl Budget

Crawl budget is the number of URLs Googlebot can and wants to crawl on a site within a given timeframe. It combines the crawl capacity limit, set by server health, with crawl demand, set by the site's popularity and staleness.

Rendering (JavaScript SEO)

Rendering is the stage where Google executes a page's JavaScript in a headless Chromium browser to see the final content. Pages that depend on client-side JavaScript are indexed from this rendered output, not the raw HTML.

Orphan Page

An orphan page is a page with no internal links pointing to it from anywhere else on the site. Crawlers that discover URLs by following links rarely find orphan pages, so they are crawled late, infrequently, or never.

XML Sitemap

An XML sitemap is a machine-readable file listing the URLs a site wants search engines to crawl, with optional metadata like last-modified dates. It is a discovery aid: listing a URL invites crawling but does not guarantee indexing.

Internal Linking

Internal linking is the practice of linking one page on a domain to another page on the same domain. Internal links are Googlebot's primary discovery channel, and they distribute authority and relevance signals across the site.

Crawl Rate

Crawl rate is the number of requests per second Googlebot makes to a site while crawling it. Google sets the rate automatically from server health signals; the manual crawl-rate limiter in Search Console was retired in January 2024.

HTTP Status Codes (for SEO)

HTTP status codes are the three-digit responses a server returns for every request. Google reads them as indexing instructions: 2xx pages can be indexed, 3xx transfer signals to the target, 404/410 drop URLs, and 5xx slows crawling.

Pronto para indexar os teus URLs?

Começa com 100 créditos grátis. Sem cartão de crédito.