SaaS, project management tools
A Mid-Size B2B SaaS Company
Google was crawling thousands of low-value URLs every week while ignoring the product pages that mattered most. Fixing the crawl pattern, not the content, moved the needle.
Key takeaways
- Noindex is not a crawl-budget fix. It stops indexing, not crawling.
- Search Console's aggregated crawl stats hid the actual problem. Raw log access made the difference between guessing and knowing.
The challenge
The site had roughly 40,000 indexed URLs, most of them internal search results and filtered category pages with no real content of their own. New product pages sometimes took three to four weeks to get crawled at all.
The team had already tried adding more internal links to new pages. It didn't help, because the crawl budget was being spent elsewhere before Google ever reached them.
The research
I pulled 30 days of server logs and cross-referenced every URL Googlebot requested against the actual sitemap. Over 60 percent of Googlebot requests in that window went to faceted navigation URLs that were already marked noindex, which meant Google was still spending crawl budget on pages it was never going to show anyone.
I also checked Search Console's crawl stats report against the raw logs, since the two did not fully agree, and the raw logs turned out to be the more reliable picture.
The solution
- Blocked the faceted navigation patterns in robots.txt instead of relying on noindex alone, since noindex still costs a crawl request.
- Rebuilt the XML sitemap to only include canonical, indexable URLs, removed roughly 28,000 low-value URLs from it.
- Added a change frequency signal to the sitemap for high-priority product pages so new releases got flagged as worth revisiting.
The result
| Metric | Before | After |
|---|---|---|
| Time to first crawl for new pages | 3–4 weeks | 2–4 days |
| Googlebot requests to low-value URLs | ~60% of crawl | ~9% of crawl |
| Indexed product pages | 340 | 410 |
New product pages started appearing in search within days instead of weeks. Organic traffic to the product section grew over the following quarter, though I want to be careful here: several product launches happened in that same window, so I can't cleanly attribute all of the growth to the crawl fix alone.
What I'd do differently
Noindex is not a crawl-budget fix. It stops indexing, not crawling. I'd lead with robots.txt exclusions earlier next time instead of treating noindex as sufficient on its own.
Search Console's aggregated crawl stats hid the actual problem. Raw log access made the difference between guessing and knowing.
Get notes like this before they're published.
Subscribe