Most audits are a crawler export with a logo on top. A useful one explains why Google is or isn't crawling, rendering, indexing and ranking your pages, traces every problem back to the template that causes it, and ends with fixes ranked by impact.
A technical audit pays for itself when something structural is wrong and nobody can say what. The usual triggers we hear:
It is not worth buying for a ten-page brochure site with no indexing problems. Run our free Marketing Analytics Auditor on it first; if that comes back clean on SEO, performance and tagging, spend the money on content instead.
Eight technical layers, each with its own primary evidence, and a written deliverable for every one.
| Layer | What we check | Primary evidence |
|---|---|---|
| Crawl access | robots.txt rules, status codes, redirect chains, soft 404s, crawl traps | Crawl + logs |
| Indexation | Indexed vs. submitted URLs, stray noindex, sitemap hygiene | Search Console |
| Canonicalization | Declared vs. Google-selected canonicals, parameter and protocol duplicates | URL Inspection + crawl |
| Rendering | Raw HTML vs. rendered DOM, client-only links and content, blocked resources | JS vs. non-JS crawl diff |
| Core Web Vitals | LCP, INP and CLS per template, field vs. lab, third-party script weight | CrUX + live trace |
| Structured data | Validity, entity consistency, markup that matches visible content | Rich Results Test + crawl |
| Internal linking | Orphans, click depth, links pointing at non-canonical URLs | Crawl graph |
| Server logs | What verified Googlebot actually requests and what it gets back | Access logs |
What lands in your inbox at the end:
Search Console (full user is enough), GA4, read access to the CMS or repository, and 30–90 days of CDN or server access logs. We export 16 months of Search Console performance data before touching anything so there's a clean before-picture.
One raw-HTML crawl, one rendered crawl. Screaming Frog's free version stops at 500 URLs and leaves out JavaScript rendering, so larger sites are crawled with a licensed copy or with Semrush or Ahrefs site audits depending on what you already pay for.
Verified Googlebot requests grouped by template, status code and response time. Crawl waste and slow templates fall out of this step.
Performance traces on the highest-value templates to name the LCP element, long tasks and third-party payloads, then cross-checked against field data.
Every issue is tied to the component, theme file or route that produces it, then scored on impact, effort and risk. Quick wins are separated from projects that need a sprint.
A working session with whoever will implement. If you want us to ship the fixes, they arrive as reviewable code changes. After deploy we re-crawl and use Search Console's validation flows to confirm.
noindex or password rule that survived launch on a subset of templates.lastmod on every URL, every day.The last two are not hypothetical. We found them on ahmeego.com.
The most common crawl mistake is using robots.txt to hide pages. Google is explicit that robots.txt
is not a mechanism for keeping a web page out of Google: a disallowed URL can still be indexed if other pages link to it, it just shows up without a description. If a
page must stay out of the index it needs noindex or authentication, and it must stay crawlable so
Google can see the noindex. We also look for parameter explosions from filters, internal search and
calendars, because every one of those URLs competes for the same crawl attention as your money pages.
We reconcile three lists: what your sitemaps submit, what the crawl finds, and what Search Console says is
indexed. Gaps between them are where the real problems live. Sitemaps get checked against Google's own limits of
50,000 URLs or 50MB uncompressed per file. The same page notes that Google ignores priority and changefreq, and Google only
trusts lastmod when it is consistently accurate. A plugin that rewrites every date daily teaches
Google to disregard the field. Google also retired the old sitemap "ping" endpoint in 2023, so submission now
happens through Search Console and the Sitemap: line in robots.txt.
Google ranks its canonical signals publicly: redirects and rel="canonical" are strong signals,
sitemap inclusion is a weak one, and they work best when they agree (Google's canonicalization documentation). Audits regularly turn up canonicals that point at redirected or 404 URLs, sitemaps listing one version while
the page declares another, and JavaScript that rewrites the canonical after load — a pattern Google specifically
warns against. We compare your declared canonical with the one Google selected in URL Inspection for a sample
from every template.
We crawl the site twice, once reading raw HTML and once rendering JavaScript as a smartphone crawler, then diff
the two. Product grids, pagination and navigation that only exist after client-side rendering show up
immediately. Links must be real <a href> elements; click handlers on a div are
invisible to a crawler. We also check that CSS and script files Google needs aren't disallowed, because blocking
resources a page depends on makes it harder for Google to understand.
The targets are fixed and public: Largest Contentful Paint within 2.5 seconds, Interaction to Next Paint of 200 milliseconds or less, and Cumulative Layout Shift of 0.1 or less, measured at the 75th percentile of page loads, split by mobile and desktop (web.dev, Web Vitals). INP replaced First Input Delay as a Core Web Vital in 2024, and lab tools like Lighthouse cannot measure it because there is no real user input, so we read field data first and use Total Blocking Time only as a lab proxy. For each template we identify the actual LCP element and the scripts holding the main thread.
We validate markup, but the bigger issues are usually consistency and honesty: Organization, Person, Product or LocalBusiness entities that contradict each other across templates, or markup describing content that isn't visible on the page. We also set expectations. Since August 2023, Google shows FAQ rich results only for well-known, authoritative government and health websites and has deprecated HowTo rich results, so we don't sell FAQ schema as a click-through play.
The crawl graph shows orphaned pages, pages buried six clicks deep, and internal links pointing at redirected or non-canonical URLs. Google recommends linking to the canonical URL consistently; every internal link to a duplicate is a mixed signal you chose to send. We also check that hub pages actually link to every child they're supposed to represent.
Logs are the only place you see what Googlebot really requested, how often, and what status it received. Before
counting anything we verify the bot, because plenty of scrapers claim to be Googlebot. Google's documented
method is a reverse DNS lookup that must resolve to a Google-owned domain such as googlebot.com,
followed by a forward lookup that returns the original IP (Google: verify requests from Google crawlers). Then we measure crawl share by template: if half of verified Googlebot hits land on filter URLs, that's your
crawl budget story in one number.
Templates are technical, and so is the risk they create. Google's spam policies define scaled content abuse and doorway abuse in terms of how pages are produced and whether they help users. If a large share of your URLs are near-identical variants, we flag it in the audit even though no crawler will call it an error. We learned that one the hard way on our own site, documented below.
Numbers we hold findings against. Each comes from the platform's own documentation or from public web-scale data, not from a tool vendor's scoring model.
| Limit or benchmark | Value | Source |
|---|---|---|
| robots.txt file size Google reads | 500 KiB; anything after that is ignored | Google robots.txt spec |
| How long Google caches robots.txt | Generally up to 24 hours, longer if refreshes fail | Same page |
| HTML Googlebot fetches per file | Up to 15MB | What is Googlebot |
| Redirect hops Googlebot follows | Up to 10, but Google advises going straight to the final URL | Site moves with URL changes |
| Sitemap file limits | 50,000 URLs or 50MB uncompressed | Build and submit a sitemap |
| When crawl budget is worth managing | Roughly 1M+ pages changing weekly, or 10,000+ changing daily | Crawl budget management |
| "Good" Core Web Vitals | LCP ≤ 2.5 s, INP ≤ 200 ms, CLS ≤ 0.1 at the 75th percentile | web.dev |
| Share of sites passing Core Web Vitals, 2025 | 48% on mobile, 56% on desktop | HTTP Archive Web Almanac 2025 |
We don't publish a single audit price because a 300-page site and a 3-million-URL catalog are different jobs. You get a scoped flat fee after a free first look at Search Console and a sample crawl. What moves the quote: URL count, number of distinct templates, whether JavaScript rendering is involved, and whether logs are available. As with everything on our pricing page, there is no percentage-of-spend fee and no long contract.
Implementation is quoted separately so you can hand the plan to your own developers. If you want ongoing work afterwards, the audit rolls into the broader SEO program.
John Williams runs the technical side: crawl and log analysis, rendering, structured data and the code changes. He built ahmeego.com's own Cloudflare stack, and before founding Ahmeego he worked at Brainlabs, OuterBox and Seer Interactive. Sandeep Muley handles the measurement layer and reporting, and Kristy Morgan keeps the fix plan moving with your team.
The audit stack. Where a client already licenses a crawler, we use theirs rather than adding another subscription.
After Google's March 2026 spam and core updates, we mapped every section of our city-page template against Google's documented standards, admitted it was templated swap content, and rebuilt it with real data.
Read the write-up →Live browser traces, a policy review of our programmatic pages, 326KB of analytics JavaScript moved off the critical path, a duplicate cache header removed, render-blocking fonts fixed — shipped in one session.
Read the audit →100+ checks across 13 categories, including SEO, security headers, accessibility, consent mode and PageSpeed / Core Web Vitals. It's the first thing we run on a new site.
Run it on your site →Our open-source audit framework connects to Search Console, GA4, GTM, PageSpeed and Merchant Center so the evidence for each finding is pulled from the source, not pasted in.
See the engine →John on what changes when answers come from AI assistants, and why crawlable, well-structured pages still sit underneath all of it.
Watch the video →Send the domain and what changed recently. We reply within one business day with what we'd look at first.
The primary sources, standards, research, and tools we rely on for this work. Every link was checked on 2026-10-11. We aren't affiliated with these publishers unless noted.