Technical SEO audit: what a real one covers, and what you walk away with

Most audits are a crawler export with a logo on top. A useful one explains why Google is or isn't crawling, rendering, indexing and ranking your pages, traces every problem back to the template that causes it, and ends with fixes ranked by impact.

TL;DR

Who this audit is for

A technical audit pays for itself when something structural is wrong and nobody can say what. The usual triggers we hear:

It is not worth buying for a ten-page brochure site with no indexing problems. Run our free Marketing Analytics Auditor on it first; if that comes back clean on SEO, performance and tagging, spend the money on content instead.

What's included

Eight technical layers, each with its own primary evidence, and a written deliverable for every one.

Layer What we check Primary evidence
Crawl access robots.txt rules, status codes, redirect chains, soft 404s, crawl traps Crawl + logs
Indexation Indexed vs. submitted URLs, stray noindex, sitemap hygiene Search Console
Canonicalization Declared vs. Google-selected canonicals, parameter and protocol duplicates URL Inspection + crawl
Rendering Raw HTML vs. rendered DOM, client-only links and content, blocked resources JS vs. non-JS crawl diff
Core Web Vitals LCP, INP and CLS per template, field vs. lab, third-party script weight CrUX + live trace
Structured data Validity, entity consistency, markup that matches visible content Rich Results Test + crawl
Internal linking Orphans, click depth, links pointing at non-canonical URLs Crawl graph
Server logs What verified Googlebot actually requests and what it gets back Access logs

What lands in your inbox at the end:

How the audit runs, step by step

Access and baseline

Search Console (full user is enough), GA4, read access to the CMS or repository, and 30–90 days of CDN or server access logs. We export 16 months of Search Console performance data before touching anything so there's a clean before-picture.

Two crawls and a diff

One raw-HTML crawl, one rendered crawl. Screaming Frog's free version stops at 500 URLs and leaves out JavaScript rendering, so larger sites are crawled with a licensed copy or with Semrush or Ahrefs site audits depending on what you already pay for.

Log analysis

Verified Googlebot requests grouped by template, status code and response time. Crawl waste and slow templates fall out of this step.

Live browser traces

Performance traces on the highest-value templates to name the LCP element, long tasks and third-party payloads, then cross-checked against field data.

Template mapping and prioritization

Every issue is tied to the component, theme file or route that produces it, then scored on impact, effort and risk. Quick wins are separated from projects that need a sprint.

Readout, fixes, re-check

A working session with whoever will implement. If you want us to ship the fixes, they arrive as reviewable code changes. After deploy we re-crawl and use Search Console's validation flows to confirm.

Problems we find again and again

The last two are not hypothetical. We found them on ahmeego.com.

The eight layers in detail

Crawl access

The most common crawl mistake is using robots.txt to hide pages. Google is explicit that robots.txt is not a mechanism for keeping a web page out of Google: a disallowed URL can still be indexed if other pages link to it, it just shows up without a description. If a page must stay out of the index it needs noindex or authentication, and it must stay crawlable so Google can see the noindex. We also look for parameter explosions from filters, internal search and calendars, because every one of those URLs competes for the same crawl attention as your money pages.

Indexation

We reconcile three lists: what your sitemaps submit, what the crawl finds, and what Search Console says is indexed. Gaps between them are where the real problems live. Sitemaps get checked against Google's own limits of 50,000 URLs or 50MB uncompressed per file. The same page notes that Google ignores priority and changefreq, and Google only trusts lastmod when it is consistently accurate. A plugin that rewrites every date daily teaches Google to disregard the field. Google also retired the old sitemap "ping" endpoint in 2023, so submission now happens through Search Console and the Sitemap: line in robots.txt.

Canonicalization

Google ranks its canonical signals publicly: redirects and rel="canonical" are strong signals, sitemap inclusion is a weak one, and they work best when they agree (Google's canonicalization documentation). Audits regularly turn up canonicals that point at redirected or 404 URLs, sitemaps listing one version while the page declares another, and JavaScript that rewrites the canonical after load — a pattern Google specifically warns against. We compare your declared canonical with the one Google selected in URL Inspection for a sample from every template.

Rendering

We crawl the site twice, once reading raw HTML and once rendering JavaScript as a smartphone crawler, then diff the two. Product grids, pagination and navigation that only exist after client-side rendering show up immediately. Links must be real <a href> elements; click handlers on a div are invisible to a crawler. We also check that CSS and script files Google needs aren't disallowed, because blocking resources a page depends on makes it harder for Google to understand.

Core Web Vitals

The targets are fixed and public: Largest Contentful Paint within 2.5 seconds, Interaction to Next Paint of 200 milliseconds or less, and Cumulative Layout Shift of 0.1 or less, measured at the 75th percentile of page loads, split by mobile and desktop (web.dev, Web Vitals). INP replaced First Input Delay as a Core Web Vital in 2024, and lab tools like Lighthouse cannot measure it because there is no real user input, so we read field data first and use Total Blocking Time only as a lab proxy. For each template we identify the actual LCP element and the scripts holding the main thread.

Structured data

We validate markup, but the bigger issues are usually consistency and honesty: Organization, Person, Product or LocalBusiness entities that contradict each other across templates, or markup describing content that isn't visible on the page. We also set expectations. Since August 2023, Google shows FAQ rich results only for well-known, authoritative government and health websites and has deprecated HowTo rich results, so we don't sell FAQ schema as a click-through play.

Internal linking

The crawl graph shows orphaned pages, pages buried six clicks deep, and internal links pointing at redirected or non-canonical URLs. Google recommends linking to the canonical URL consistently; every internal link to a duplicate is a mixed signal you chose to send. We also check that hub pages actually link to every child they're supposed to represent.

Server logs

Logs are the only place you see what Googlebot really requested, how often, and what status it received. Before counting anything we verify the bot, because plenty of scrapers claim to be Googlebot. Google's documented method is a reverse DNS lookup that must resolve to a Google-owned domain such as googlebot.com, followed by a forward lookup that returns the original IP (Google: verify requests from Google crawlers). Then we measure crawl share by template: if half of verified Googlebot hits land on filter URLs, that's your crawl budget story in one number.

The check nobody puts in a technical audit: policy risk

Templates are technical, and so is the risk they create. Google's spam policies define scaled content abuse and doorway abuse in terms of how pages are produced and whether they help users. If a large share of your URLs are near-identical variants, we flag it in the audit even though no crawler will call it an error. We learned that one the hard way on our own site, documented below.

Limits and benchmarks to know

Numbers we hold findings against. Each comes from the platform's own documentation or from public web-scale data, not from a tool vendor's scoring model.

Limit or benchmark Value Source
robots.txt file size Google reads 500 KiB; anything after that is ignored Google robots.txt spec
How long Google caches robots.txt Generally up to 24 hours, longer if refreshes fail Same page
HTML Googlebot fetches per file Up to 15MB What is Googlebot
Redirect hops Googlebot follows Up to 10, but Google advises going straight to the final URL Site moves with URL changes
Sitemap file limits 50,000 URLs or 50MB uncompressed Build and submit a sitemap
When crawl budget is worth managing Roughly 1M+ pages changing weekly, or 10,000+ changing daily Crawl budget management
"Good" Core Web Vitals LCP ≤ 2.5 s, INP ≤ 200 ms, CLS ≤ 0.1 at the 75th percentile web.dev
Share of sites passing Core Web Vitals, 2025 48% on mobile, 56% on desktop HTTP Archive Web Almanac 2025

Pricing and engagement

We don't publish a single audit price because a 300-page site and a 3-million-URL catalog are different jobs. You get a scoped flat fee after a free first look at Search Console and a sample crawl. What moves the quote: URL count, number of distinct templates, whether JavaScript rendering is involved, and whether logs are available. As with everything on our pricing page, there is no percentage-of-spend fee and no long contract.

Implementation is quoted separately so you can hand the plan to your own developers. If you want ongoing work afterwards, the audit rolls into the broader SEO program.

John Williams runs the technical side: crawl and log analysis, rendering, structured data and the code changes. He built ahmeego.com's own Cloudflare stack, and before founding Ahmeego he worked at Brainlabs, OuterBox and Seer Interactive. Sandeep Muley handles the measurement layer and reporting, and Kristy Morgan keeps the fix plan moving with your team.

Platforms we work in

The audit stack. Where a client already licenses a crawler, we use theirs rather than adding another subscription.

Google Search Console logoGoogle Search Console Bing Webmaster Tools logoBing Webmaster Tools Screaming Frog SEO Spider Semrush logoSemrush Ahrefs Lighthouse / PageSpeed Insights logoLighthouse / PageSpeed Insights Google Analytics 4 logoGoogle Analytics 4 Google Tag Manager logoGoogle Tag Manager Cloudflare logoCloudflare

Proof you can check: we audit our own site in public

Case study · ahmeego.com

Rebuilding 2.8 million programmatic pages

After Google's March 2026 spam and core updates, we mapped every section of our city-page template against Google's documented standards, admitted it was templated swap content, and rebuilt it with real data.

Read the write-up →
Case study · ahmeego.com

The audit an AI agent ran on itself

Live browser traces, a policy review of our programmatic pages, 326KB of analytics JavaScript moved off the critical path, a duplicate cache header removed, render-blocking fonts fixed — shipped in one session.

Read the audit →
Free tool

Marketing Analytics Auditor

100+ checks across 13 categories, including SEO, security headers, accessibility, consent mode and PageSpeed / Core Web Vitals. It's the first thing we run on a new site.

Run it on your site →
Free tool

250-point Audit Engine

Our open-source audit framework connects to Search Console, GA4, GTM, PageSpeed and Merchant Center so the evidence for each finding is pulled from the source, not pasted in.

See the engine →
Video

Navigating the shift from SEO to LLMs

John on what changes when answers come from AI assistants, and why crawlable, well-structured pages still sit underneath all of it.

Watch the video →

Frequently asked

How long does a technical SEO audit take?
The free first look comes back within 48 hours, as described on our SEO page. The full audit's timeline is set in the scope and depends mostly on site size and how quickly server logs can be exported; log access is usually the slowest step, not the crawl.
Do you really need my server logs?
You'll still get a solid audit without them, but logs are the only direct evidence of what Googlebot requests and what it receives. For large catalogs and sites with indexing problems they are the most valuable input. We verify Googlebot by reverse and forward DNS before counting any request.
My traffic dropped after a core update. Will a technical audit fix it?
Possibly, but be careful with anyone who promises that. Core updates mostly reassess content quality and usefulness. We check the technical layers and the policy layer together — duplicated templates, thin programmatic pages, doorway patterns — because recovering usually means changing what the pages say, not only how they are served.
Will you fix the issues or just report them?
Either. Many clients hand the plan to in-house developers. If you want us to implement, we ship fixes as reviewable code changes and re-check them after deploy.
Does FAQ or HowTo schema still help?
Not as a rich-result play for most sites. Since August 2023 Google limits FAQ rich results to well-known government and health sites and has deprecated HowTo rich results. Accurate structured data still helps machines understand a page, so we keep markup that matches visible content and drop the rest.
Can you audit React, Next.js or other JavaScript sites?
Yes. Rendering is one of the eight layers. We diff raw HTML against the rendered DOM, check the canonical is set in the HTML source rather than rewritten by script, and confirm links are real anchor elements Google can follow.

Request a technical SEO audit

Send the domain and what changed recently. We reply within one business day with what we'd look at first.

John, Kristy, or Sandeep will reply. One of the three of us will respond personally within 1 business day. No SDR queue.
We respond within 1 business day. No spam, ever. Read our privacy notice.

Top 25 references

The primary sources, standards, research, and tools we rely on for this work. Every link was checked on 2026-10-11. We aren't affiliated with these publishers unless noted.

Official documentation

  1. Introduction to robots.txt — Google Search Central
    Google's own statement that robots.txt manages crawling and is not a way to keep a page out of the index.
  2. Build and submit a sitemap — Google Search Central
    The 50MB / 50,000-URL file limits and the note that Google ignores priority and changefreq values.
  3. What is URL canonicalization / Consolidate duplicate URLs — Google Search Central
    Ranks redirects and rel=canonical as strong signals and sitemap inclusion as a weak one.
  4. Understand the JavaScript SEO basics — Google Search Central
    How Google crawls, renders and indexes JavaScript pages, including guidance on canonicals and links set by script.
  5. Spam policies for Google web search — Google Search Central
    The definitions of scaled content abuse and doorway abuse we use to flag risky page templates.
  6. Verify requests from Google crawlers and fetchers — Google Crawling Infrastructure
    The reverse-then-forward DNS method for confirming a log entry really came from Googlebot.
  7. Crawl budget management — Google Crawling Infrastructure
    Tells you whether crawl budget is your problem at all: written for 1M+ page sites or 10,000+ pages that change daily.
  8. Page indexing report — Search Console Help
    What each 'not indexed' reason means, from 'Crawled - currently not indexed' to Google-selected canonical conflicts.
  9. URL Inspection tool — Search Console Help
    How to compare the canonical you declared with the one Google selected, and test the live rendered page.
  10. Bing Webmaster Guidelines — Microsoft Bing
    Bing's crawl, indexing and quality rules, which matter because Bing's index feeds several AI assistants.
  11. Web Vitals — web.dev (Google Chrome)
    The official LCP, INP and CLS thresholds and the 75th-percentile rule for assessing a page.
  12. Interaction to Next Paint (INP) — web.dev (Google Chrome)
    How INP is measured across every interaction on a page and why it replaced First Input Delay.
  13. Optimize Largest Contentful Paint — web.dev (Google Chrome)
    Breaks LCP into four sub-parts so you can tell a slow server from a late-discovered hero image.

Standards & policy

  1. RFC 9309: Robots Exclusion Protocol — IETF
    The formal robots.txt standard, useful when a crawler and a CMS disagree on how a rule should match.
  2. Sitemaps XML format — sitemaps.org
    The protocol itself: required tags, index files, encoding and the W3C date format for lastmod.
  3. RFC 9110: HTTP Semantics — IETF
    The definitive meaning of 301, 308, 404, 410 and 5xx responses that every crawl finding depends on.

Research & studies

  1. Web Almanac 2025: SEO — HTTP Archive
    Web-scale data on robots.txt, canonicals, raw versus rendered tags and structured data across millions of sites.
  2. Web Almanac 2025: Performance — HTTP Archive
    Reports that 48% of mobile sites had good Core Web Vitals in 2025 and about 16-17% lazy-load their LCP image.
  3. Chrome UX Report (CrUX) documentation — Chrome for Developers
    How the real-user field dataset behind Search Console's Core Web Vitals report is collected and queried.

Leading tools

  1. Screaming Frog SEO Spider — Screaming Frog
    The desktop crawler we use for raw and rendered crawls; the free version stops at 500 URLs and omits JavaScript rendering.
  2. Screaming Frog Log File Analyser — Screaming Frog
    Imports access logs, verifies search bots and reports crawl frequency and status codes by URL.
  3. Rich Results Test — Google
    Shows which rich result types Google can read from a URL or code snippet and flags invalid properties.
  4. Lighthouse overview — Chrome for Developers
    Lab audits for performance and SEO basics; useful for diagnosis, but lab scores are not field Core Web Vitals.

Expert guides

  1. The Beginner's Guide to Technical SEO — Ahrefs
    A clear walkthrough of crawling, indexing and site-structure fundamentals for teams new to technical SEO.
  2. Technical SEO (Beginner's Guide to SEO, chapter 5) — Moz
    Long-standing primer on how rendering, page speed and structured data fit together, good for briefing developers.
AI disclosure: This page was drafted with AI assistance and edited by a human. Third-party facts link to the official source they came from (checked 2026-10-11); platform names and logos belong to their owners and do not imply a partnership or endorsement.