|

Tiempo de lectura

11 min

What a sitemap is and what it is for

A sitemap is a list of the URLs you want known. The XML is for Google. The HTML is a page for people. It is not a Google Maps map and it is not a floor plan. This page separates the two files and says which one to submit to Search Console. The XML does not index a URL just by listing it.

Two things with the same name

An XML sitemap is a file. An HTML sitemap is a page. If your question is "where am I on the map," this is not the page. If your question is why Google does not know about a new URL, it is.

The XML is the one you submit to Search Console. The HTML is optional. A ten-page site that is well linked can live without an HTML index. Almost none should live without knowing what their XML is and which URLs it contains.

What the XML does, and what it does not

It hands Google a list. The bot can discover URLs from links, from history, and from this list. On a small, well-linked site, the sitemap is not the reason you get indexed. On a new site, or after publishing a block of pages, it shortens the time until Google learns they exist.

It does not do these things, even when they are sold that way:

  • It does not index. A URL in the sitemap and on noindex stays out. The noindex wins.

  • It does not rank. Being on the list is not being on page one.

  • It does not clean errors. If the URL redirects or returns 404, putting it in the sitemap asks Google to look at a problem.

  • It does not replace internal links. A page that only lives in the sitemap, and that nobody links, is an orphan with a mention in a file.

What it should include, and what should stay out

  1. The canonical URLs you want in the index: services, contact, current articles, pages that sell.

  2. Out: the ones that redirect. The list should carry the destination, not the origin of the 301 redirect.

  3. Out: the ones with noindex: thank-you pages, drafts, internal search results.

  4. Out: parameter variants that are not a different page.

  5. Out: test URLs and the old domain if you already migrated.

If the file and the index contradict each other, what Google finds by crawling wins. The sitemap does not undo a canonical or a block. That priority is in what technical SEO is.

Where you look

Most platforms generate it themselves, at a path like /sitemap.xml. Open it in the browser. You should see URLs on your domain, not the example template. If you see pages you deleted months ago, the generator has not updated or you are looking at an old static file.

In Search Console, the sitemaps report says whether Google could read it and how many URLs it discovered. "Discovered" does not mean "indexed." The pages report is the one that says whether they were stored. Confusing the two numbers is the reading that causes the most unnecessary content plans.

In robots.txt you can leave the line that points at the file. It is a help, not an obligation if you already submitted it in Search Console. Do not block the sitemap with robots.txt itself.

The HTML, so you do not mix them

An HTML sitemap is a page of links, often in the footer, of the "site map" kind. It helps when navigation is narrow and a person cannot reach a section. Google can follow those links too, as it follows any other page. It is not uploaded as a file. It does not use XML tags. And it is not used to hide, in a huge list, what you did not want to link in the menu: if it should not be indexed, do not link it there either.

After a migration or a deletion

Regenerate the list. Remove the URLs that already redirect. Check that the new ones are in. Look at Search Console again in case the file it reads is still the old one. The context of that change is in migrating without losing rankings.

If the site has few links and pages even you cannot find, the sitemap is not the project. The architecture is. The file then only documents a site a person cannot walk through either. When the list is clean, making each URL deserve to be on it is SEO, not generating another XML.

A check before you submit it

Open the XML and read ten URLs at random, not only the first. Ask whether you would want that URL to appear in Google. If the answer is no, take it off the list before you submit it. Check there are no test domains, no http mixed with https, and not the same page with and without a trailing slash if one of the two redirects.

After you submit it, note the date. A week later, see in Search Console how many of those URLs are actually indexed. If the file has forty and the index has twelve, do not generate another sitemap. Read the exclusion reason for the ones that are missing. It is almost always a noindex, a canonical, or a page Google treats as a duplicate, not a file that "was not processed."

If the platform regenerates the file on its own, confirm it does not put back what you just removed. A generator that lists everything published, including thank-you pages and drafts, undoes the work every night. In that case the fix is to exclude those templates in the generator, not to delete one URL by hand every Monday.

XML sitemap and HTML sitemap

One file is for the search engine and one page is for people. The table says who each one is for, and what happens if you treat them backwards.


XML sitemap

HTML sitemap

Who it is for

Google and other search engines

A person who cannot find a page

Where it lives

A file, often /sitemap.xml

A normal URL on the site, linked in the footer

What it achieves

The bot learns URLs it may not have reached by links

A person can navigate. It does not index anything by itself

What it does not do

It does not force indexing or remove a noindex

It does not replace the XML and it is not submitted as a file

Conclusion

Publish an XML that lists only the URLs you want in the index, submit it in Search Console, and do not use it as proof that you are already indexed. The HTML, if you need it, is another page: an index for people, not a file for the bot.

Frequently asked questions about sitemaps

What is a sitemap?

It is a list of addresses on your website. In XML, you give that list to Google so it can discover URLs. In HTML, it is a page that helps a person find sections. It is not a city map and it is not the Google Business Profile. Listing a URL also does not force Google to index it.

What is an XML sitemap for?

It tells the search engine which URLs matter to you, especially if the site is new, if it has few internal links, or if you have just published many pages. It does not index. Google can ignore a URL on the list if it is noindex, if it redirects, or if it does not deserve to be in the index.

Do you have to submit the sitemap to Google?

Yes, once, in Search Console, in the sitemaps report. You can also declare it in robots.txt. Submitting it is not an indexing button. It is how Google finds the list. After that you have to see whether those URLs are actually indexed.

Is a sitemap a location map?

No. This query gets mixed up with maps of physical places. A website sitemap does not show a street or a floor plan. If you want to appear in Google Maps, the work is the business profile, not the XML file. They are two different projects.

Need a marketing team?

Let’s talk and grow your business

BG Image

Let's build something remarkable

Ready for your next project

+

You

15-minute call

Pick the best time for you

Vector
Vector
Element Image
Otros blogs