A planting calendar for search work done in-house
Packet 03 · Sitemaps and crawlingExample sowing: May–June
Sitemaps and crawling, worked through
Google mainly finds pages by following links from pages it already knows, so a small site where every page is linked may not need a sitemap at all. Google adds: “However, in most cases, your site will benefit from having a sitemap.” Where one helps, it is a plain file listing the pages you want in Google’s results, and this page builds one for an example business, step by step.
How Google finds a page
Google Search works in three stages, and not every page makes it through each one.
- Crawling. There is no central register of every web page, so Google keeps looking for new and updated ones. It finds them by following links from pages it knows, or from a list you submit, called a sitemap. Its crawler, Googlebot, then downloads what is on them.
- Indexing. Google analyses the text, images and video on the page and stores the information in its index, a large database.
- Serving. When someone searches, Google returns information relevant to the query.
A sitemap only touches the first stage. Google calls submitting one “merely a hint”: there is no promise that Google will fetch the file, or crawl the addresses listed in it.
Do you need one?
Google’s sitemap overview gives two lists.
| You might need a sitemap if | You might not need one if |
|---|---|
| Your site is large, so it is harder to make sure every page is linked from another. | Your site is small: about 500 pages or fewer that you want in search results. |
| Your site is new and has few links to it from other sites. | Googlebot can reach every important page by following links from the home page. |
| You have a lot of video and images, or appear in Google News. | You don’t have many media files or news pages you want shown in search. |
If your site runs on a content management system such as WordPress, Wix or Blogger, Google says it has likely already made a sitemap available, and you don’t have to do anything.
The worked example
Dotto’s example business, invented for this page
A bakery has a website at https://www.example.com/ (Google’s own placeholder domain in its documentation). The site has these addresses:
| Address | In the sitemap? | Google’s reason |
|---|---|---|
/ (home) | Yes | Include the URLs you want to see in Google’s search results. |
/bread/ | Yes | Same rule: a page you want found. |
/bread/sourdough.html | Yes | Same rule. |
/cakes/ | Yes | Same rule. |
/visit.html | Yes | Same rule. |
/bread/?sort=price | No | The same content under another address. Choose the URL you prefer and list only that one. |
/order/thanks.html | No | A page the bakery doesn’t want in search results, so it doesn’t belong in a list of pages it does. |
Step 2: count
Five pages go in. Google’s line for a “small” site is about 500 pages or fewer, so by Google’s measure this site is small. If every page can be reached by links from the home page, Google’s Sitemaps report help says a site like this probably doesn’t need a sitemap: requesting indexing of the home page is enough. The bakery builds one anyway, below, so you can see the shape.
Step 3: write the file
Google accepts XML, RSS or Atom, and plain text sitemaps, and has no preference between them. For fewer than a few dozen URLs, Google says you may be able to write one by hand in a text editor. Here is the bakery’s, in XML, following the shape of Google’s own basic example:
sitemap.xml (example, in the format of Google’s basic XML sitemap)
<?xml version="1.0" encoding="UTF-8"?>
<!-- Dotto's example: one url entry for each page the bakery wants found -->
<urlset xmlns="http://www.sitemaps.org/schemas/sitemap/0.9">
<!-- the home page -->
<url>
<loc>https://www.example.com/</loc>
<lastmod>2026-09-30</lastmod>
</url>
<url>
<loc>https://www.example.com/bread/</loc>
</url>
<url>
<loc>https://www.example.com/bread/sourdough.html</loc>
<lastmod>2026-08-14</lastmod>
</url>
<url>
<loc>https://www.example.com/cakes/</loc>
</url>
<url>
<loc>https://www.example.com/visit.html</loc>
</url>
</urlset>
What each choice in that file follows:
- Full addresses. Use fully qualified, absolute URLs; Google crawls them exactly as listed.
- Dates only where true. Google uses
lastmodonly if it is consistently and verifiably accurate, and it should reflect the last significant update, such as a change to the main content, not a new copyright year. The example dates are invented; leave the field out rather than guess. - No priority or change frequency. Google ignores the
priorityandchangefreqvalues, so the example leaves them out. - Order doesn’t matter. Google says the order of URLs in a sitemap doesn’t matter to it.
- Encoding and place. The file must be UTF-8. Google recommends posting it at the site root, where it can cover every file on the site.
- Size. As at October 2026, Google’s page sets a limit of 50MB uncompressed or 50,000 URLs per sitemap; beyond that, split it into several.
The same list as a text sitemap is one URL per line and nothing else, in a file with a .txt extension:
sitemap.txt (example)
https://www.example.com/
https://www.example.com/bread/
https://www.example.com/bread/sourdough.html
https://www.example.com/cakes/
https://www.example.com/visit.html
Step 4: tell Google where it is
Google lists several ways to make a sitemap available; two suit a small site.
- In robots.txt. Add a line anywhere in the file naming the sitemap, and Google finds it the next time it crawls robots.txt:
Sitemap: https://www.example.com/sitemap.xml - In the Sitemaps report. You need owner permission on the Search Console property. First check that Googlebot can reach the file: run a live URL inspection on its address and look for a Page fetch of “Successful”. Then paste the address into “Add a new sitemap” and select Submit.
Step 5: read the status
The Sitemaps report shows one of three results for the latest fetch: Success, Has errors (fetched, but with one or more errors; URLs that parsed cleanly are still queued for crawling) or Couldn’t fetch. Google lists reasons it may not fetch a sitemap, including a robots.txt rule blocking it, an unresolved manual action, or a wrong address.
The report only lists sitemaps submitted through it or the API, not ones found through robots.txt. Google says you can still submit a sitemap it already knows about, to track its success and errors.
Crawling: the gate and the sign
Two tools look alike and do different jobs. A robots.txt rule stops crawling; a noindex rule stops indexing. Google says not to use robots.txt to keep a page out of search, but to use noindex or a login.
The two can cancel each other out. For noindex to work, the page must not be blocked by robots.txt, because a crawler that can’t reach the page never sees the rule, and the page can still appear in results if other pages link to it. And keep CSS and JavaScript open to Google: if they are hidden, Google might not be able to understand your pages.
All six packets
Packet 01Jan–Feb
What an SEO manager looks after
Packet 02Mar–Apr
Search Console in plain words
Packet 03May–Jun
Sitemaps and crawling
Packet 04Jul–Aug
Helpful content, as a checklist
Packet 05Sep–Oct
Fixing what Google can’t index
Packet 06Nov–Dec
What structured data does