Semrush helps you:

  • Do keyword research
  • Audit your local listings
  • Perform competitor analyses
  • Manage social media accounts
  • And much more!

Backlinko readers get:

A 14-day trial for premium features. 55+ tools.
Free access for core features.

Newsletter Sign Up

Backlinko readers get
access for 14 days. 55+ tools.

Technical SEO: The Definitive Guide

yongi-barnard

Written by Yongi Barnard

Technical SEO Guide – Featured image

Every SEO investment you make — content, links, authority — depends on one thing working first.

Technical SEO.

Before search engines can rank your pages, they need to find them, crawl them, render them, and trust what they say.

Now, AI search systems rely on those same foundations.

If your content is blocked, hard to interpret, or technically unreliable, it’s less likely to show up in AI-generated answers.

This guide covers the technical SEO fundamentals you need to get right. Plus, the AI-specific layers that now belong in the same checklist.

Free resource: Download our FREE technical SEO scorecard. Stop guessing what’s broken on your website and know exactly where to focus first.

Technical SEO Fundamentals

Technical SEO is really about one thing: Removing the obstacles that stop machines from finding, accessing, understanding, and using your content.

That includes search engines. And increasingly, AI systems.

What Is Technical SEO?

Technical SEO is the process of improving your website so search engines and AI systems can crawl, render, index, and reuse your content.

Key components of technical SEO include:

  • Website architecture
  • XML sitemaps
  • Robots.txt
  • Internal linking
  • Core Web Vitals
  • Mobile usability
  • Structured data
  • Site performance
  • Content extractability
  • And more
Key components of technical SEO

Why Is Technical SEO Important?

Technical SEO is important because it directly affects whether your pages get discovered, how they rank, and how users experience them.

Technical SEO…

  • Protects the user experience: Fast-loading, secure, mobile-optimized pages keep readers around (and satisfy Google’s Core Web Vitals standards)
  • Consolidates ranking signals: Fixing duplicate content and canonical issues keeps your ranking signals from splitting across competing URLs
  • Improves your odds of being ranked highly and cited: Clean, consistent structure makes it easier for AI systems to extract your content and reuse it in answers

That’s not to say your technical SEO has to be perfect. It doesn’t.

But the easier you make it for search engines and AI crawlers, the better your visibility in both.

Strong technical SEO

Search engines and AI search systems work differently. But they share the same first requirement:

They must be able to process your web content before they can use it.

From processing to visibility

Most technical SEO problems are a breakdown at a stage in that process.

Finding where the breakdown happens helps you choose the right fix.

Site Structure and Internal Links

Site structure is how your pages are organized and interlinked.

A clear structure helps people and machines find your pages, move through your site, and understand how each page relates to the others.

Get this right, and the rest of technical SEO falls into place:

  • Crawlers can discover important pages more efficiently
  • Internal links help distribute authority and context
  • Search engines and AI search systems get clearer signals about what your site covers and which pages matter most

This is why website architecture is one of the first things I look at in any technical SEO campaign, even before crawling reports and indexing errors.

Build a Hierarchical Website Architecture

A hierarchical architecture organizes your site into clear levels, with pages nested under logical parent sections.

Your homepage sits at the top.

Under that, your main categories. Each category can branch into subcategories, which can then branch into sub-subcategories as needed.

Website architecture example

The ideal number of levels depends on the size and complexity of your site.

That hierarchy helps machines understand how topics are organized and which subject areas are most important.

To check your site’s hierarchy, use the Semrush Site Audit tool.

Go to “SEO” > “Site Audit” > “Crawled Pages” > “Site Structure.”

Site Audit – Backlinko – Crawled pages – Site structure

Check that there’s a clear path from the homepage to your main sections and supporting pages.

Backlinko’s structure is a clean example.

Site Audit – Backlinko – Site structure

The main sections — Blog, Hub, Templates — sit directly under the homepage and then branch into subcategories.

You can apply this same pattern to different types of sites.

For example:

  • Ecommerce: Homepage > Products > Running Shoes > Trail Running Shoes
  • SEO agency: Homepage > Services > Technical SEO Audits
  • SaaS: Homepage > Features > Specific Feature Page

Optimize Click Depth

Click depth is the number of clicks it takes to get from your homepage to another page on your site.

Your most important pages should ideally be reachable in three clicks or fewer.

Click depth

That’s a guideline, not a hard rule — bigger sites may naturally run deeper. What matters is keeping priority pages close to the top.

Click depth affects how easily crawlers find and prioritize pages.

Ones closer to the homepage tend to get crawled more often and pick up stronger internal linking signals.

It matters for AI search, too. A page buried deep in your site is harder for crawlers to reach and process — which makes it less likely to surface in search results or AI answers.

To check click depth, run a crawl in Screaming Frog and sort by the “Crawl Depth” column.

Screaming Frog – Report – Crawl Depth

Pay close attention to pages that target competitive keywords or explain core products and services.

These should not be buried too far from the homepage.

If they are, link to them where they fit naturally: the homepage, parent hub pages, relevant category pages, or related supporting content.

Strengthen Internal Links

Internal links are any links from one page on your site to another.

These include links in your main navigation, body content, footer, sidebar, breadcrumbs, and related-content modules.

Backlinko – Homepage links

A strong internal linking structure makes important pages easier for users and machines to find.

It also helps prevent orphan pages: pages with no internal links pointing to them.

(These are harder for crawlers to discover.)

Orphan pages have no internal links

But as a starting point:

  • Link high-traffic posts to relevant commercial pages
  • Link commercial pages with supporting guides and proof content
  • Make sure every important page has at least one link pointing to it
  • Link between related pages within the same topic cluster to reinforce topical relevance

An audit tool can help you assess your structure by showing how many internal links point to each page.

This makes it easy to see which pages are well supported and which ones are orphans.

Screaming Frog – Report – Inlinks

Use Breadcrumbs to Confirm Page Relationships

Breadcrumbs help search engines understand your site structure by showing how pages relate to one another.

A consistent breadcrumb path reinforces where a page belongs and how it connects to the rest of the site.

Semrush – Breadcrumbs

They also improve technical SEO because they:

  • Create crawlable internal links
  • Connect deep pages to key categories or hubs
  • Reduce orphan-page risk
  • Help users navigate the site

Clear breadcrumbs leave no room for ambiguity.
Like this Sephora product page’s breadcrumb trail: Makeup > Cheek > Blush.

Sephora – Breadcrumbs

The page’s place in the site architecture is immediately clear to crawlers and users.
Makeup is the broad category. Cheek is the subcategory. Blush is the product category.
Two best practices for breadcrumbs:

First, make them visible and crawlable.

They must appear on the page and use standard HTML links, not JavaScript-only breadcrumbs.

Breadcrumbs as standard HTML links

Second, keep breadcrumbs consistent.

Each page should occupy one spot in your hierarchy, and its breadcrumb should reflect that same spot everywhere the page appears.

Make URLs Descriptive and Consistent

URLs help search engines and AI systems understand what a page is about before they read and interpret the content.

That means URLs play a role in context and machine understanding.

Before a search engine fully processes a page’s content, it encounters the URL.

The URL can provide early clues about the page’s topic and relationship to other pages on the site.

Good URL example

Compare these:

code icon
yourwebsite.com/technical-seo/xml-sitemaps
yourwebsite.com/p?id=4827&cat=12

The first URL signals the topic and site structure.

It shows the page is about XML sitemaps and belongs under technical SEO.

The second URL provides less context.

URL structure

Machines can still process it, but they need additional signals to do so — such as titles, headings, and internal links.

Multiply that across a site with hundreds or thousands of pages, and you create unnecessary interpretation friction.

Use Canonical Tags to Clarify the Preferred URL

A canonical tag is a snippet of code in a page’s <head> that tells search systems which URL you’d prefer as the primary version.

Atlassian – Canonical URL

It points crawlers to a single, clear source rather than any competing duplicates.

Say you have three different URLs with the same content:

code icon
/womens-running-shoes
/womens-running-shoes/
/womens-running-shoes?ref=nav

Search engines and LLMs may treat those as separate pages.

That can waste crawl resources and potentially surface a URL you don’t want in the SERPs.

You fix that with a canonical link element on every duplicate or alternate URL like this:

<link rel=”canonical” href=”https://www.yoursite.com/womens-running-shoes/” />

That signals to machines: “Treat this URL as the preferred version of the page.”

Shopify uses canonicals

Keep in mind it’s a strong hint, not a command. Search engines usually respect it, but can pick a different URL if other signals disagree.

Also add that same tag to the primary page — a self-referencing canonical — to reinforce that it’s the version you want indexed.

Crawling, Rendering, and Indexing

This chapter is about making it easy for search engines and AI crawlers to find, render, and index the pages that matter.

I’ll show you how to:

  • Spot crawl barriers
  • Fix access issues
  • Help crawlers reach important pages deeper in your site

Configure Robots.txt and Crawler Access

Robots.txt tells crawlers which parts of your site they should crawl and which to avoid.

Use it to manage crawler access to duplicate, low-value, or non-search areas, such as internal search results and checkout pages.

Here’s an example of Walmart’s robots.txt:

Walmart – Robots.txt – Disallow gift registry

This robots.txt allows crawlers to reach pages that matter for search, like product pages, reviews, and a store finder.

But it blocks crawlers from certain areas, like the gift registry page.

Walmart – Robots.txt – Disallow gift registry

Two things to never block:

Pages you want to show up in search AND the JavaScript, CSS, or images Google needs to render your pages.

To make sure you aren’t blocking anything important, perform a manual check by typing yoursite.com/robots.txt into your browser.

Then look at each “Disallow” line.

For each one, ask: “Do I want crawlers in there?”

Alternatively, you can also use Google Search Console.

Go to “Indexing” > “Pages” > “Not Indexed.”

Then scroll to the “Why pages aren’t indexed report.”

Look for “Blocked by robots.txt” and review those URLs.

GSC – Indexing – Blocked by robots

Make sure none of them are pages you want crawled and ranked.

Note: You may have heard of LLMs.txt. It is a proposed file that tells AI systems what to prioritize when crawling your site. Google Search Central says it’s not necessary for its AI search surfaces. But these files may become useful as AI agents interact with websites more directly.

Build a Reliable XML Sitemap

An XML sitemap tells crawlers which URLs you want them to discover.

You can usually find it at “yoursite.com/sitemap.xml.”

(Though some sites use sitemap indexes or custom locations.)

Backlinko – XML Sitemap

Your sitemap should only include URLs you want discovered and indexed.

Meaning, they are:

  • Canonical: Preferred version
  • Indexable: Not blocked
  • Live: Returns a 200 status code
  • Useful: Page with search value

Don’t include redirects, broken pages, duplicate URLs, parameter URLs, or pages you don’t want indexed.

Backlinko’s free sitemap generator can help you create a sitemap quickly.

Backlinko – Sitemap generator

To make sure your sitemap is working correctly, check its status in Google Search Console.

Go to “Indexing” > “Sitemaps” to confirm it has been submitted and processed. And check for any errors Google flagged.

GSC – Submitted sitemaps

Then use Semrush for a deeper check.

Go to “Site Audit” > “Issues” and search for “sitemap” to surface problems like incorrect or orphaned pages.

Site Audit – Backlinko – Issues – Sitemap

Don’t Hide Content Behind JavaScript

Many AI crawlers don’t fully render JavaScript.

Google search can render it, but the extra rendering step can delay indexing.

So if important content loads only through JavaScript, AI systems may not see it — and search engines may index it less reliably.

To check whether content depends on JavaScript, disable JavaScript in Chrome and reload your priority pages.

Go to:

Settings” > “Privacy and security” > “Site settings” > “JavaScript.”

Reload the page, then see what disappears.

Here’s an example from Crutchfield, an electronics retailer, where important product text does not appear when JavaScript is disabled.

JavaScript enabled vs disabled

There’s no one-size-fits-all fix, especially on larger sites.

Work with your developer to serve important content server-side, so machines get it whether or not they render JavaScript.

Find and Fix Important Pages That Aren’t Indexed

If a page isn’t discoverable, accessible, or indexed by the systems that matter, it’s unlikely to rank or appear in AI search.

Not every unindexed page is a problem.

But revenue-driving pages — especially product pages, service pages, and landing pages — should be indexable and accessible to major search systems.

Use Google Search Console as a diagnostic starting point.

It only shows how Google indexes your site, but its Page Indexing report can reveal technical issues that may affect broader search and AI visibility.

Go to Google Search Console > Indexing > “Pages.”

GSC – Indexing Pages – Overview

Scroll to “Why pages aren’t indexed” and check for common technical issues, such as:

  • noindex tags
  • Blocked URLs
  • Broken pages
  • Canonical issues

Fix the issue based on the report.

Then use the URL Inspection tool to request that Google recrawl the page.

Site Performance and Reliability

Search engines and AI systems regularly crawl your pages to keep their information up to date.

A slow or unreliable site gets in the way of that.

Here’s what to address to make your site easy to reach, fast to load, and reliable.

Improve Page Speed and Core Web Vitals

Core Web Vitals (CWV) are Google metrics that measure real user experience on a page.

The three Core Web Vitals are:

  • Largest Contentful Paint (LCP): How quickly the main content becomes visible. Aim for 2.5 seconds or less.
  • Interaction to Next Paint (INP): How quickly the page responds to user interaction. Aim for 200 milliseconds or less.
  • Cumulative Layout Shift (CLS): How much the layout unexpectedly moves. Aim for 0.1 or less.
Core Web Vitals

Depending on your traffic, you can monitor Core Web Vitals in Google Search Console under:

Experience > “Core Web Vitals.”

GSC – Core Web Vitals

Optimize the Mobile Experience

Google predominantly uses the mobile version of your site for indexing and ranking.

If your mobile experience is slow or broken, it doesn’t matter how good the desktop version is.

That’s why mobile SEO is a core part of technical SEO.

To see how mobile readers experience your page, you can use Google’s Lighthouse.

Chrome – Lighthouse

Lighthouse will identify the elements that ruin the mobile experience, including:

  • Text that’s too small to read comfortably
  • Buttons and links that are too close together
  • Content extending beyond the viewport

Use HTTPS to Make Your Site Secure

Site security is a foundational requirement for technical SEO.

Users and crawlers should be able to access your pages through a secure connection, without browser warnings or inconsistent HTTP/HTTPS signals.

HTTPS

For most sites, the core security requirement is HTTPS.

That sounds basic.

After all, more than 90% of mobile websites already use HTTPS, as the Web Almanac reports.

HTTPS usage

The problem is when the main domain is secure while other parts aren’t.

This can happen when:

  • An old subdomain still loads over HTTP
  • A migrated section is missing a valid SSL certificate
  • A certificate has expired
  • HTTPS pages load images, scripts, fonts, or iframes over HTTP
  • Internal links, canonicals, redirects, or XML sitemaps still point to HTTP URLs

When this happens, crawlers may find duplicate HTTP URLs, inconsistent signals, or resources that don’t load cleanly.

So don’t just check whether the homepage has a padlock.

Backlinko – Secure connection

Check that the secure version works everywhere.

You can check this in Screaming Frog by crawling all versions of your site, like:

  • www.yoursite.com
  • blog.yoursite.com
  • app.yoursite.com
Screaming Frog – Check multiple URLs

Make sure each one redirects cleanly to the HTTPS version.

Keep Your Site Stable

When crawlers request an important page, your site should deliver it reliably.

If search engines repeatedly run into timeouts, server errors, DNS problems, or missing pages, they have a harder time crawling your site efficiently.

This doesn’t necessarily mean a full-site outage.

Your homepage might load fine while product pages time out.

Or the site might have DNS issues, server connectivity problems, or robots.txt fetch failures that only affect crawlers.

To check these issues, go to “Settings” > “Crawl stats” in Google Search Console.

GSC – Settings – Crawl stats

For URL-level issues, go to “Indexing” > “Pages” and look for server errors, 404s, and soft 404s.

These errors aren’t always urgent.

A few old 404s from deleted pages may be fine. What matters is the pattern.

GSC – Soft 404 indexing error

If you see a spike in errors, look at what changed around the same time.

  • Was there a site launch?
  • A migration?
  • A CMS update?
  • A template change?

Patterns usually tell you where the real problem is.

For ongoing monitoring, use a tool like UptimeRobot that monitors your pages and alerts you if your webpage goes down or suddenly slows down.

UptimeRobot status

Monitor your most important URLs, such as service pages, checkout pages, lead forms, and booking flows.

That way, you get alerted when something breaks instead of discovering the issue later in Search Console.

Semantic HTML and Structured Data

Semantic HTML and structured data give machines clearer context about your pages.

Semantic HTML clarifies the page’s structure, while structured data turns key details into standardized, machine-readable information.

Use Semantic HTML

Semantic HTML tags describe the meaning and structure of each part of a page.

They make it easier for search engines and AI systems to identify your headings, navigation, main content, and other key sections.

Compare these two code blocks:

Non-semantic vs semantic HTML code

The non-semantic version uses <div> and <span> tags and doesn’t indicate what each section represents.

But in the semantic version, each part is clearly labeled <header>, <nav>, <main>, or <footer>.

This reduces ambiguity, especially on large or template-driven sites where the same navigation, product cards, and footer elements appear across thousands of pages.

To check your pages, run them through the W3C Markup Validator.

W3C – Markup Validation Service

Or for a quick first pass, you can also use an AI tool.

Copy the HTML of the page you’re auditing, paste it into an LLM, and ask it to audit the semantic structure.

It will give you a pretty good summary of the page’s semantic HTML status.

ChatGPT – HTML audit

The results give you a useful starting point and a concrete basis for a conversation with your developer.

Use this prompt in any AI tool:

Audit this HTML for semantic structure.

Check:

1. Whether the page uses proper landmarks like header, nav, main, article, section, aside, and footer.
2. Whether there is one clear H1.
3. Whether headings follow a logical hierarchy.
4. Whether links use <a href=””> and buttons use <button>.
5. Whether important content is trapped in generic divs when a semantic element would be better.
6. Whether the structure would still make sense without CSS.

Return:

– What is good
– What is unclear
– What should be changed
– Example corrected HTML

Side note: Semantic HTML is not a ranking factor. And your code does not need to be perfect for search engines and AI systems to understand your page. But the cleaner your structure is, the less your site forces machines to guess.

Implement Structured Data

Structured data is standardized, machine-readable labels you add to your code.

It helps search engines and AI systems understand what type of page they’re processing.

It also identifies elements on the page, including products, organizations, FAQs, services, authors, prices, and ratings.

This gives machines clearer context about your content.

It can also make your information easier to interpret and use in search features and AI answers.

Recipe snippet example

Most structured data uses Schema.org vocabulary written as JSON-LD inside a <script> tag.

Like this:

Schema Markup Code

The schema you implement depends on what’s on the page.

Product pages get product markup, including name, description, image, price, availability, and reviews.

JDSports – Product schema

A blog post gets markup such as the headline, author, publish date, and featured image.

Whatever schema you use, one rule applies:

The markup must match what users can see on the page.

If an item is out of stock on the page, the structured data should not say otherwise. Otherwise, search engines may ignore the structured data or treat it as misleading or spammy.

(And as quality standards tighten for AI search, accurate markup will likely become even more important.)

The Verge – Google AI search spam policy

To add schema markup to a page, start with your CMS.

These usually have built-in schema options or plugins that handle common markup types for you.

For WordPress, Yoast and Rank Math are common options. For Shopify, many themes and apps can automatically generate product schema.

And for pages where details rarely change, you can add JSON-LD manually.

Backlinko’s free Schema Markup Generator can help you create the code.

Backlinko – Schema markup generator

Clarify Entity Identity

Entity identity is how search engines and AI systems recognize your brand as a specific “thing” on the web.

It helps them understand what you do and who you serve. Plus, how your site connects to related products, people, services, and topics.

Many areas of SEO contribute to entity identity.

But in technical SEO, the focus is on making signals easy for machines to parse and connect.

A big part of this is schema.

Google Developers – Structured data – Local business

For example, “Organization” schema defines your brand, website, logo, contact details, and official profiles.

You can also use “sameAs” links to connect your website to official external profiles.

As Patagonia does on its homepage.

Patagonia – Organization schema

Other ways to strengthen entity identity:

  • Keep identity signals consistent: Use the same brand name, logo, contact details, author names, product names, and official profile links across your site
  • Create a clear identity hub: Use your homepage and about page to explain who you are, what you do, who you serve, and what topics you’re known for
  • Show relationships with internal links: Link authors to their articles, services to case studies, products to categories, and guides to related topic pages

To check that you’ve declared your site’s entity identity on your homepage or about page, install the Detailed SEO Extension.

Go to your homepage, click the plugin, and open the Schema tab.

Detailed SEO Extension – Able Carry – Schema

Related reading: Want to learn more about building entity identity signals? Read this article on AI Search Trust Signals.

Technical SEO Maintenance and Monitoring

Technical SEO is not a one-time setup.

Your site changes over time, and those changes can create new issues that affect how search engines and AI systems access and understand your content.

Regular monitoring helps you catch problems early so that your site stays fast and secure.

It also helps keep you visible in search and AI systems.

Analyze Log Files

Log files show how search and AI crawlers interact with your site.

Unlike a site crawl, which shows what crawlers could access, log files reveal what they actually visited.

For example, log files can show whether:

  • Search engine crawlers reached new category pages after a launch
  • Important product pages are being crawled regularly
  • Bots are hitting 500 errors, redirect chains, or blocked URLs
  • AI crawlers like GPTBot, ClaudeBot, and PerplexityBot are reaching pages that matter
Log file

Use a log file analyzer to surface patterns and issues quickly.

Tools like Screaming Frog and Semrush can help you identify crawl activity, errors, and access problems before they affect rankings or traffic.

Log File Analyzer – Googlebot Activity

Audit Your Technical SEO Health

Technical SEO issues are easier to fix when you catch them early.

To do this, use a site audit tool to scan your site on a regular schedule.

Tools like Screaming Frog and Semrush can automatically flag problems for you.

These tools can also help you find issues like:

  • Broken links and broken pages
  • Redirect chains
  • Duplicate content
  • Thin pages
  • Missing metadata

Run a full audit quarterly as a baseline.

Site Audit – Backlinko – Overview

You should also run one after major site changes, including:

  • A site migration
  • A sudden drop in organic traffic
  • A major launch or redesign
  • A CMS or template change
  • A significant algorithm update

This gives you a repeatable way to monitor the health of technical SEO.

Technical SEO Is Your AI Search Foundation

AI search hasn’t made technical SEO less important. If anything, it’s raised the stakes on getting the basics right.

Search engines and AI search systems need to access, process, and understand your content before they can use it.

If your important pages are blocked, buried, or slow, they get less consideration from both.

You don’t need to do everything now.

Start with your most important money pages.

Use our free Technical SEO Scorecard to identify your biggest gaps.

When you’re ready for a full site audit, read our guide on the best SEO audit tools.

Backlinko is owned by Semrush. We’re still obsessed with bringing you world-class SEO insights, backed by hands-on experience. Unless otherwise noted, this content was written by either an employee or paid contractor of Semrush Inc.