This is a concept website · Enquire about this domain

SEO Manager

A planting calendar for search work done in-house

Back to the planting calendar

Packet 03 · Sitemaps and crawlingExample sowing: May–June

Sitemaps and crawling, worked through

Google mainly finds pages by following links from pages it already knows, so a small site where every page is linked may not need a sitemap at all. Google adds: “However, in most cases, your site will benefit from having a sitemap.” Where one helps, it is a plain file listing the pages you want in Google’s results, and this page builds one for an example business, step by step.

How Google finds a page

Google Search works in three stages, and not every page makes it through each one.

  1. Crawling. There is no central register of every web page, so Google keeps looking for new and updated ones. It finds them by following links from pages it knows, or from a list you submit, called a sitemap. Its crawler, Googlebot, then downloads what is on them.
  2. Indexing. Google analyses the text, images and video on the page and stores the information in its index, a large database.
  3. Serving. When someone searches, Google returns information relevant to the query.

A sitemap only touches the first stage. Google calls submitting one “merely a hint”: there is no promise that Google will fetch the file, or crawl the addresses listed in it.

Do you need one?

Google’s sitemap overview gives two lists.

Google’s own lists, side by side
You might need a sitemap ifYou might not need one if
Your site is large, so it is harder to make sure every page is linked from another. Your site is small: about 500 pages or fewer that you want in search results.
Your site is new and has few links to it from other sites. Googlebot can reach every important page by following links from the home page.
You have a lot of video and images, or appear in Google News. You don’t have many media files or news pages you want shown in search.

If your site runs on a content management system such as WordPress, Wix or Blogger, Google says it has likely already made a sitemap available, and you don’t have to do anything.

The worked example

Dotto’s example business, invented for this page

A bakery has a website at https://www.example.com/ (Google’s own placeholder domain in its documentation). The site has these addresses:

Step 1: list every address, then decide which belong in the sitemap
AddressIn the sitemap?Google’s reason
/ (home)YesInclude the URLs you want to see in Google’s search results.
/bread/YesSame rule: a page you want found.
/bread/sourdough.htmlYesSame rule.
/cakes/YesSame rule.
/visit.htmlYesSame rule.
/bread/?sort=priceNoThe same content under another address. Choose the URL you prefer and list only that one.
/order/thanks.htmlNoA page the bakery doesn’t want in search results, so it doesn’t belong in a list of pages it does.

Step 2: count

Five pages go in. Google’s line for a “small” site is about 500 pages or fewer, so by Google’s measure this site is small. If every page can be reached by links from the home page, Google’s Sitemaps report help says a site like this probably doesn’t need a sitemap: requesting indexing of the home page is enough. The bakery builds one anyway, below, so you can see the shape.

Step 3: write the file

Google accepts XML, RSS or Atom, and plain text sitemaps, and has no preference between them. For fewer than a few dozen URLs, Google says you may be able to write one by hand in a text editor. Here is the bakery’s, in XML, following the shape of Google’s own basic example:

sitemap.xml (example, in the format of Google’s basic XML sitemap)

<?xml version="1.0" encoding="UTF-8"?>
<!-- Dotto's example: one url entry for each page the bakery wants found -->
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
  <!-- the home page -->
  <url>
    <loc>https://www.example.com/</loc>
    <lastmod>2026-09-30</lastmod>
  </url>
  <url>
    <loc>https://www.example.com/bread/</loc>
  </url>
  <url>
    <loc>https://www.example.com/bread/sourdough.html</loc>
    <lastmod>2026-08-14</lastmod>
  </url>
  <url>
    <loc>https://www.example.com/cakes/</loc>
  </url>
  <url>
    <loc>https://www.example.com/visit.html</loc>
  </url>
</urlset>

What each choice in that file follows:

  • Full addresses. Use fully qualified, absolute URLs; Google crawls them exactly as listed.
  • Dates only where true. Google uses lastmod only if it is consistently and verifiably accurate, and it should reflect the last significant update, such as a change to the main content, not a new copyright year. The example dates are invented; leave the field out rather than guess.
  • No priority or change frequency. Google ignores the priority and changefreq values, so the example leaves them out.
  • Order doesn’t matter. Google says the order of URLs in a sitemap doesn’t matter to it.
  • Encoding and place. The file must be UTF-8. Google recommends posting it at the site root, where it can cover every file on the site.
  • Size. As at October 2026, Google’s page sets a limit of 50MB uncompressed or 50,000 URLs per sitemap; beyond that, split it into several.

The same list as a text sitemap is one URL per line and nothing else, in a file with a .txt extension:

sitemap.txt (example)

https://www.example.com/
https://www.example.com/bread/
https://www.example.com/bread/sourdough.html
https://www.example.com/cakes/
https://www.example.com/visit.html

Step 4: tell Google where it is

Google lists several ways to make a sitemap available; two suit a small site.

  1. In robots.txt. Add a line anywhere in the file naming the sitemap, and Google finds it the next time it crawls robots.txt:
    Sitemap: https://www.example.com/sitemap.xml
  2. In the Sitemaps report. You need owner permission on the Search Console property. First check that Googlebot can reach the file: run a live URL inspection on its address and look for a Page fetch of “Successful”. Then paste the address into “Add a new sitemap” and select Submit.

Step 5: read the status

The Sitemaps report shows one of three results for the latest fetch: Success, Has errors (fetched, but with one or more errors; URLs that parsed cleanly are still queued for crawling) or Couldn’t fetch. Google lists reasons it may not fetch a sitemap, including a robots.txt rule blocking it, an unresolved manual action, or a wrong address.

The report only lists sitemaps submitted through it or the API, not ones found through robots.txt. Google says you can still submit a sitemap it already knows about, to track its success and errors.

Crawling: the gate and the sign

Two tools look alike and do different jobs. A robots.txt rule stops crawling; a noindex rule stops indexing. Google says not to use robots.txt to keep a page out of search, but to use noindex or a login.

The two can cancel each other out. For noindex to work, the page must not be blocked by robots.txt, because a crawler that can’t reach the page never sees the rule, and the page can still appear in results if other pages link to it. And keep CSS and JavaScript open to Google: if they are hidden, Google might not be able to understand your pages.

All six packets