Handbook / Module 2 / Lesson 4

Soft 404s, Redirect Errors, and Blocked by Robots.txt

Diagnose and resolve Soft 404 algorithms, fix complex redirect loops and chains, and master the nuances of robots.txt indexing exclusions.

Intermediate 18 min read #Soft 404 #Redirects #Robots.txt #HTTP Status Codes

What is a Soft 404?

A Soft 404 is not an official HTTP status code. It is an algorithmic classification applied by Google when a page returns an HTTP status of 200 OK, but its content suggests to Google’s machine learning classifiers that the page is actually missing, empty, or dead.

┌────────────────────────────────────────────────────────────────────────┐
│  SERVER SAYS:      HTTP/1.1 200 OK                                     │
│  CONTENT SAYS:     "Sorry, no products found in this category."        │
│                    OR "Page does not exist - redirected to home."      │
│                                                                        │
│  GOOGLE CLASSIFIES: Soft 404 (Excluded from Index)                     │
└────────────────────────────────────────────────────────────────────────┘

What Triggers a Soft 404?

  1. Category Pages with Zero Inventory: E-commerce search results or product category filters that currently return zero products.
  2. Generic Error Templates with HTTP 200: A CMS that renders a friendly “Oops! We could not find what you were looking for” graphic while sending a 200 OK header instead of a true 404 Not Found or 410 Gone.
  3. Blank Single Page Apps (SPAs): A client-side routing bug where API data fails to load, leaving only header, footer, and a blank content container.
  4. Aggressive Homepage Redirects: Catch-all redirects sending deleted product URLs back to https://example.com/.
A popular brand deleted 8,000 obsolete products and decided to redirect all 8,000 URLs to their homepage (`/`) using 301 redirects, thinking "we won't lose any link equity!"

The Result: Google classified all 8,000 redirects as Soft 404s. Google does not transfer PageRank or ranking signals through catch-all homepage redirects when the destination page is completely irrelevant to the original deleted product.


Diagnosing Redirect Errors

The Redirect error status appears when Googlebot is unable to reach the destination of a redirect.

Common Redirect Faults:

  1. Redirect Loops: URL A redirects to URL B, which redirects back to URL A.
  2. Excessive Redirect Chains: More than 5 consecutive redirect hops (A -> B -> C -> D -> E -> F). Googlebot typically halts crawling after 5 hops.
  3. Empty or Malformed Target: The server responds with Location: (blank string) or relative path errors like Location: //https://....
  4. Protocol Mismatch: Redirecting across insecure or untrusted SSL configurations.

“Blocked by robots.txt” vs. “Indexed, though blocked”

Developers are frequently surprised when a URL they explicitly disallowed in robots.txt still appears in Google Search results!

┌────────────────────────────────────────────────────────────────────────┐
│               THE ROBOTS.TXT CRAWL VS. INDEX DISTINCTION               │
│                                                                        │
│  robots.txt prevents CRAWLING, but it does NOT prevent INDEXING!       │
└────────────────────────────────────────────────────────────────────────┘
  • If an external website or internal page links to https://example.com/secret/, Google knows the URL exists.
  • If robots.txt contains Disallow: /secret/, Googlebot obeys and never fetches the page contents.
  • Because Google cannot crawl the page, it cannot read <meta name="robots" content="noindex">!
  • Google will still index the raw URL string without a snippet, showing the message: “No information is available for this page.”
To properly remove a page from Google's index: 1. Ensure the page is **ALLOWED** in `robots.txt` so Googlebot can fetch it. 2. Add `` in the HTML ``, or send the HTTP header `X-Robots-Tag: noindex`. 3. Once Google crawls the page and confirms the `noindex` directive, you can disallow it in `robots.txt` if necessary.

Lab Challenge: Fix Soft 404s

1. In the Page Indexing report, click on **Soft 404**. 2. Examine the list of sample URLs. Are they out-of-stock products, empty categories, or custom 404 templates? 3. Configure your web server to return a genuine `404 Not Found` or `410 Gone` HTTP status code for permanently discontinued content.