Crawling is the process in which a bot discovers URLs and downloads their content. Only then can a search engine process the page and decide whether to index it.
What is crawling?
Googlebot visits known URLs, downloads the HTML and follows links. How often it comes back to new or changed pages varies with the site, the importance of the URL and how the content changes.
How bots discover pages
The main sources are internal and external links, XML sitemaps and previously known addresses. An orphaned page with no internal link is harder to discover, and the architecture says nothing about its importance.
A sitemap helps bots find URLs but doesn't guarantee crawling or indexing. It should list the canonical, indexable pages you want in search.
What slows crawling down
- blocks in
robots.txt, - server errors and long outages,
- redirect loops or chains,
- endless combinations of filters and parameters,
- links without a real
hrefattribute, - content that only loads after a user interaction.
A successful download still doesn't guarantee indexing or a ranking. Crawling is the first technical prerequisite of SEO. The results come later.