Handbook / Module 1 / Lesson 1

Deconstructing the URL Inspection Tool

Master every diagnostic data point in the URL Inspection tool: crawl verdict, discovery paths, crawl bot identity, indexing allowance, and canonical declarations.

Intermediate 18 min read #URL Inspection #Diagnostics #Canonical #Googlebot

The Heart of Search Console Debugging

The URL Inspection Tool (accessible via the top global search bar in Search Console) provides a direct probe into Google’s database for any specific URL within your verified property.

When you enter a URL, GSC queries the live Google Index and returns the exact state of that page as of its last recorded crawl.


Anatomy of the URL Inspection Verdict

Google Search Console URL Inspection Tool Interface Figure 1.1: The URL Inspection dashboard demonstrating the ‘URL is on Google’ status, Discovery routes, crawl metadata, and canonical consensus diagnostics.

┌────────────────────────────────────────────────────────────────────────┐
│  URL is on Google / URL is not on Google                              │
│  "It can appear in Google Search results (if not subject to manual     │
│   action or removal request)"                                          │
└────────────────────────────────────────────────────────────────────────┘

The verdict summary is followed by three critical telemetry panels:

  1. Presence on Google
  2. Page Indexing (Coverage Details)
  3. Enhancements & Structured Data

Diagnostic Checklist: Key Fields & What They Mean

1. Discovery

  • Sitemaps: Lists which submitted XML sitemaps referenced this URL. If empty, Google discovered this URL via links rather than your sitemap feed.
  • Referring page: The external or internal URL Googlebot traversed to discover this link. If this says “None detected,” Google discovered it through a previously indexed version or a private link graph.

2. Crawl Telemetry

  • Last crawl: Precise UTC timestamp of Googlebot’s most recent fetch. If this date is weeks old, Google has deprioritized recrawling this page.
  • Crawled as:
    • Googlebot smartphone (Primary mobile-first indexing bot)
    • Googlebot desktop (Rare, typically reserved for resources that are restricted to desktop user agents)
  • Crawl allowed?: Checks whether your robots.txt rules permitted the crawl. (Yes vs No: blocked by robots.txt).
  • Page fetch: The HTTP transport result:
    • Successful (HTTP 200)
    • Failed: Not found (404)
    • Failed: Server error (5xx)
    • Failed: Redirect error
  • Indexing allowed?: Checks whether <meta name="robots" content="noindex"> or an X-Robots-Tag: noindex header was returned.

3. Canonicalization

This is the most critical technical diagnostic in the report:

  • User-declared canonical: The exact URL your code placed in <link rel="canonical" href="...">.
  • Google-selected canonical: The URL Google chose as the authoritative master copy.
When **User-declared canonical** and **Google-selected canonical** differ, Google has rejected your canonical instruction! Google treats canonical tags as *hints*, not directives. If your internal links, sitemaps, or content similarity point toward another page, Google will override your code and index the alternate URL instead.

Step-by-Step Diagnostic Workflow

[ Step 1: Input URL in Top Bar ]
               │
               ▼
   [ Check "Last Crawl" Date ]
         ├── If older than recent deploy ──> Run "Test Live URL"
         └── If recent ──> Inspect Canonical & Coverage
               │
               ▼
[ Compare User Canonical vs Google Canonical ]
         ├── Match ──> Page Canonicalized Correctly
         └── Mismatch ──> Investigate duplicate signals, redirect loops
If Google selects a different canonical than your declared URL, copy the **Google-selected canonical** and paste it back into the URL Inspection tool. Check its last crawl date and referring pages to see why Google favors that version.

Lab Challenge: Inspect Your Key Landing Page

1. Open Google Search Console and inspect your website's highest-revenue or primary product page. 2. Note the **Last crawl** timestamp and verify that **Crawled as** reports `Googlebot smartphone`. 3. Confirm that **User-declared canonical** matches **Google-selected canonical** exactly (including trailing slash and protocol).