Handbook / Module 2 / Lesson 2

Troubleshooting 'Discovered' vs 'Crawled' - Currently Not Indexed

Deconstruct the two most infamous indexing exclusions in Google Search Console, diagnose crawl budget vs quality filters, and implement step-by-step solutions.

Advanced 22 min read #Discovered Not Indexed #Crawled Not Indexed #Crawl Budget #Content Quality

The Great Indexing Dilemma

No two status messages cause more debate among developers and SEOs than:

  1. Discovered - currently not indexed
  2. Crawled - currently not indexed

While they look similar in the interface, their technical root causes exist at opposite ends of Google’s pipeline.


Detailed Comparative Diagnostic

┌───────────────────────────────────────────────┬───────────────────────────────────────────────┐
│       DISCOVERED - CURRENTLY NOT INDEXED      │         CRAWLED - CURRENTLY NOT INDEXED       │
├───────────────────────────────────────────────┼───────────────────────────────────────────────┤
│ The URL was discovered (via sitemap or link), │ Googlebot requested, downloaded, and parsed   │
│ but Googlebot has NOT yet fetched or crawled  │ the HTML/DOM, but deliberately chose NOT to   │
│ the URL due to crawl budget or server load.   │ write the page into the primary search index. │
├───────────────────────────────────────────────┼───────────────────────────────────────────────┤
│ ROOT CAUSE:                                   │ ROOT CAUSE:                                   │
│ • Server latency / host crawl capacity limits │ • Content quality / low information gain      │
│ • Excessive crawl bloat (faceted navigation)  │ • Near-duplicate content / thin templates     │
│ • Domain authority / trust threshold too low  │ • Algorithmic quality evaluation threshold    │
└───────────────────────────────────────────────┴───────────────────────────────────────────────┘

Deep Dive: “Discovered - Currently Not Indexed”

When Google marks a URL as Discovered - currently not indexed, Google’s crawler made a deliberate decision to postpone crawling.

Primary Causes & Solutions:

  1. Server Overload Signals: If your origin server experiences high TTFB (>1,000ms) or frequent HTTP 503 or 429 spikes, Googlebot reduces its crawl rate to protect your infrastructure.
    • Fix: Optimize server caching, deploy edge caching (Cloudflare, Fastly), or upgrade origin compute resources.
  2. Crawl Waste & URL Bloat: Generating thousands of parameterized URLs (?color=blue&size=m&sort=price_asc) exhausts Googlebot’s crawl budget before it ever reaches genuine content.
    • Fix: Apply robots.txt disallows or canonical tags on parameter matrices.
  3. Internal Linking Orphans: URLs submitted in an XML sitemap but receiving zero internal links across the website are deprioritized by Google’s discovery scheduler.
    • Fix: Link to newly published articles or products directly from high-authority hub pages or category nodes.
An online footwear retailer launched a faceted search filter where every combination of size, color, width, and brand generated a unique URL. Within two weeks, Search Console reported **420,000 URLs in 'Discovered - currently not indexed'**, while new brand-new product pages were waiting 3 weeks just to get their first crawl.

Remediation: The engineering team added a Disallow: /*?*filter= rule in robots.txt and purged parameter URLs from the XML sitemap. Within 14 days, Googlebot re-allocated its crawl bandwidth back to primary product pages.


Deep Dive: “Crawled - Currently Not Indexed”

When a URL enters Crawled - currently not indexed, Googlebot successfully connected to your server, executed rendering, and read your page. It simply decided that the page was not worth storing in the index.

Primary Causes & Solutions:

  1. Thin or Auto-Generated Content: Pages with only 1–2 sentences, placeholder boilerplate, or purely scraped vendor descriptions.
    • Fix: Merge thin pages into comprehensive pillar guides, or apply <meta name="robots" content="noindex"> to low-value landing pages.
  2. Near-Duplicate Content: Multiple landing pages created for slight geographic variations (e.g., “Plumber in Dallas”, “Plumber in Fort Worth”) with 95% identical text.
    • Fix: Add unique localized value, customer testimonials, specific local pricing, or canonicalize to a regional hub.
  3. Low Information Gain: If 10 other websites already provide identical explanations with higher domain authority, Google’s Helpful Content System may discard your article to save index space.
Export 100 URLs from "Crawled - currently not indexed". Ask yourself honestly: *If a searcher lands on this page, does it provide unique data, original research, or distinct functionality that does not exist elsewhere on our site or the web?* If the answer is no, prune the page or combine it with a stronger resource.

Lab Challenge: Triage Your Unindexed URLs

1. Navigate to **Page Indexing > Crawled - currently not indexed**. 2. Click **Export** to download the CSV. 3. Classify the URLs by path pattern (`/category/`, `/product/`, `/tag/`). 4. Does the majority belong to a specific dynamic template or tag archive? Formulate a pruning or consolidation strategy.