Field Data First Technical SEO Checklist for SEOs and Engineers

This technical SEO checklist gives you an ordered, audit-first sequence to find and fix the issues most likely to stop pages from being crawled, indexed, or surfaced. Start by verifying indexing, Core Web Vitals, and server errors, then move through performance, mobile, JavaScript rendering, and structured data. Work in an audit, fix, monitor loop so each change gets validated before you move to the next layer.
TL;DR:
- Use real visit data at the 75th percentile: LCP must be 2.5 seconds or less, INP 200 milliseconds or less, and CLS 0.1 or less.
- Check Search Console indexing against a rough site search, inspect key URLs, and confirm sitemaps contain only canonical, indexable pages returning 200.
- For JavaScript pages, compare raw and rendered HTML; choose server rendering for critical content if rendering delays are slowing indexing.
- For large catalogs, keep valuable filter pages indexable but block low value combinations only after confirming they offer no unique content.
- After each fix, verify changes with field data; review Core Web Vitals monthly and Search Console errors weekly during active remediation.
Table of Contents
- Crawlability and indexing checks
- Core Web Vitals and performance checklist (LCP, INP, CLS)
- Mobile friendliness and UX checks
- JavaScript SEO and rendering: how to verify what Google sees
- URL structure, redirects, and canonicalization
- Sitemaps, robots.txt, and crawl budget management
- Structured data, metadata, and snippet optimization
- Site architecture and internal linking at scale
- Security, HTTPS, and server configuration checks
- Audit tools and a step-by-step action plan
- Practical engineering notes on remediation trade-offs
- Handling duplicate content issues beyond canonical tags
- Pagination and hreflang for international and multi-language sites
- How we run a technical SEO audit: a practitioner’s note
- Running this checklist tells you what’s broken. Fixing a rendering pipeline, a CDN misconfiguration, or a faceted navigation system at scale is a different kind of work, and it’s a type of work that marketing and engineering teams often need to handle together.
- FAQ
- Sources
Crawlability and indexing checks
Before touching performance or metadata, confirm that Google can actually find and index your pages. A site that ranks well on paper but is quietly losing pages from the index will not benefit from any other fix on this list.
Start in Google Search Console’s Coverage report to see how many URLs are indexed versus excluded, and cross-check with a site: search for a rough sanity count. Neither method is exact, but a large gap between what you expect and what shows up is a signal worth chasing.
From there, use URL Inspection on a sample of important pages, including new ones and ones you suspect are struggling. It shows the actual rendered HTML and tells you whether the page is indexed, excluded, or queued, which is the closest you can get to seeing the page the way Google’s systems see it, per Google’s crawling and indexing guidance.
Next, test robots.txt directly. Common mistakes include blocking entire CSS or JS directories (which breaks rendering), blocking staging paths that later get copied into production rules, and wildcard rules that unintentionally catch important URL patterns.
Finally, validate your sitemap and check server logs for crawl activity:
- Confirm the sitemap only lists canonical, indexable, 200-status URLs, not redirects or noindexed pages
- Check that every important template (product, category, blog post) is represented in the sitemap
- Pull server logs to see which bots are hitting which URLs, and how often
- Flag any spike in 5xx server errors or DNS failures during crawl windows, since repeated server errors can cause Google to slow or pause crawling
- Compare log-confirmed crawl activity against Search Console’s crawl stats to spot gaps
A clean baseline here is what makes every later fix measurable: you need to know pages are indexable before you can judge whether a speed or markup fix moved the needle.
Core Web Vitals and performance checklist (LCP, INP, CLS)
Once indexing is confirmed, performance is the next lever with the clearest tie to both rankings and conversion. Google measures three Core Web Vitals at the 75th percentile of real visits: Largest Contentful Paint (LCP), Interaction to Next Paint (INP), and Cumulative Layout Shift (CLS).
A page needs LCP at or under 2.5 seconds, INP at or under 200 milliseconds, and CLS at or under 0.1 to reach a “Good” score at the 75th percentile, according to Google’s Core Web Vitals documentation. Scores above 4.0 seconds for LCP, 500 milliseconds for INP, or 0.25 for CLS are classified as poor.
The distinction between field data and lab data matters more than most audits treat it. Field data, pulled from the Chrome User Experience Report (CrUX) or your own real-user monitoring (RUM), reflects what actual visitors experienced on real networks and devices. Lab tools like Lighthouse or PageSpeed Insights run a single simulated session and are useful for debugging a specific page, but they will not tell you whether your real traffic is passing or failing the thresholds. Web recommends relying on field data for the actual pass or fail judgment and reserving lab tools for root-cause diagnosis.
INP deserves particular attention because it replaced First Input Delay as the responsiveness metric and catches problems that older audits missed entirely. Per web.dev’s INP optimization guide, the recommended workflow is to identify slow interactions in the field first, reproduce them in the lab by following the same user flow, then fix the underlying script work rather than guessing from a single synthetic test.
Common fixes map cleanly to each metric:
- For LCP, compress and properly size hero images, use a content delivery network, and cut server response time for the initial HTML
- For INP, break up long JavaScript tasks, defer non-critical scripts, and reduce the work done inside event handlers
- For CLS, reserve space for images and embeds with explicit width and height attributes, and avoid injecting content above existing content after load
- For all three, remove render-blocking CSS and JavaScript from the critical rendering path
Prioritize by business impact rather than by raw page count. For teams running Shopify storefronts specifically, a RUM-first performance audit can show where template-level choices are quietly adding to LCP before a redesign makes it worse.
Mobile friendliness and UX checks
Google indexes the mobile version of your site by default, which means mobile usability problems are indexing problems, not just a secondary concern. A handful of specific checks catch most of the common failures.
Confirm the page has a correct viewport meta tag, that the layout reflows properly at common mobile widths, and that tap targets (buttons, links, form fields) are large enough and spaced enough to tap without mis-hits. Cramped navigation menus and overlapping buttons are still common on sites that were designed desktop-first and adapted afterward.
Remove or resize intrusive interstitials, the full-screen pop-ups that cover content immediately on page load on mobile. A dismissible banner or a small bar is fine; a pop-up that blocks the article behind it is not. Avoid any setup that serves different content to mobile crawlers than to mobile users, since that crosses into cloaking regardless of intent.
A short checklist for this section:
- Run URL Inspection’s mobile rendering view to confirm the page renders the way a mobile visitor would see it
- Check Search Console’s mobile usability report for flagged errors like text too small to read or clickable elements too close together
- Test the actual checkout or lead form flow on a real phone, not just a browser’s device emulator
- Confirm images and videos scale correctly instead of forcing horizontal scroll
- Verify font sizes are legible without a pinch-to-zoom gesture
These checks are quick to run and catch problems that otherwise surface as high mobile bounce rates with no obvious explanation in the analytics.
JavaScript SEO and rendering: how to verify what Google sees
JavaScript-heavy sites fail audits in a specific way: the page looks fine in a browser but Google’s crawler sees something thinner or entirely different. Understanding why requires understanding how Google actually processes a JavaScript page.
Per Google’s guidance on fundamentals and its JavaScript SEO basics documentation, Google handles JS pages in phases: it crawls the raw HTML first, queues the page for rendering, then executes the JavaScript to produce the final rendered HTML before indexing it. That rendering step can be delayed, which means a change you make in JavaScript may not be reflected in the index as quickly as a static HTML change would be.
To verify what Google actually sees on a JS-heavy page, work through this order:
- Run URL Inspection and view the rendered HTML tab to confirm your content, links, and metadata appear after rendering
- Use the Rich Results Test to check whether structured data survives the rendering process intact
- Compare the raw server response against the rendered output to spot content that only appears after client-side execution
- Check that routing uses the History API rather than URL fragments, since fragment-based routing can produce URLs Google will not treat as distinct pages
- Confirm critical resources like JS bundles and APIs are not blocked by robots.txt, since a blocked resource cannot be rendered at all
For SEO-critical pages, server-side rendering or a hybrid approach (rendering key content on the server while keeping interactivity client-side) removes the rendering queue delay entirely and gives you direct control over what the crawler receives on first contact. Single-page applications without server rendering carry the highest risk; frameworks with built-in server-side rendering or incremental static regeneration carry the least, since the HTML is already complete before any JavaScript executes. A developer-focused JavaScript SEO audit walks through this verification sequence in more depth for engineering teams managing the handoff between frontend frameworks and crawlability.
Pro Tip: Diff the raw HTML response against the rendered HTML in URL Inspection before assuming a JavaScript framework migration is SEO-safe.

URL structure, redirects, and canonicalization
Canonicalization problems rarely come from one broken tag. They come from conflicting signals scattered across redirects, canonical tags, sitemaps, and internal links that each point somewhere slightly different.
Per Google’s canonicalization documentation, a canonical tag is a hint, not a directive. Google weighs rel=canonical alongside sitemap inclusion, redirect targets, and even protocol preference (HTTPS over HTTP) before choosing which URL represents a set of duplicates. That means a single correct canonical tag sitting next to an inconsistent sitemap or a redirect pointing elsewhere can still get overridden.
Treat this as signal stacking rather than a single fix:
- Set a self-referential canonical tag on every indexable page, pointing to itself
- Make sure your sitemap only contains the canonical version of each URL, never a duplicate or parameter variant
- Audit redirect chains and collapse any URL that redirects more than once before reaching its final destination
- Use 301 redirects for permanent moves and reserve 302 or 307 for genuinely temporary changes, since a permanent move marked as temporary can delay canonical consolidation
- Keep internal links pointed at the canonical URL consistently across navigation, footer, and body content
Never try to solve canonicalization with robots.txt. Blocking a duplicate URL in robots.txt prevents Google from crawling it, which means Google cannot see the canonical tag on that page at all, and the duplicate can still get indexed from external links with no way to consolidate it. Google’s own crawling and indexing guidance is explicit that rel=canonical or a redirect is the correct tool, not a disallow rule.
Sitemaps, robots.txt, and crawl budget management
Sitemaps and robots.txt are the two files that most directly shape how a search engine spends its crawl budget on your site, and both are easy to get subtly wrong at scale.
For sitemaps, split large sites into multiple sitemap files organized by section (products, categories, blog), and roll them up under a single sitemap index file once you pass a few thousand URLs. Keep the lastmod field accurate, since Google uses it as a signal for whether a URL needs to be recrawled, and a sitemap where every URL shows the same lastmod date regardless of actual changes loses its usefulness.
For robots.txt, block resources that have no SEO value, like internal search result pages, admin paths, and duplicate parameter combinations, but avoid the common over-blocking mistake of disallowing entire folders that also contain CSS, JS, or images needed for rendering.
For very large catalogs or sites with heavy dynamic content, crawl budget becomes a real constraint rather than a theoretical one:
- Use index sitemaps to help Google prioritize which sections matter most
- Keep the sitemap limited to canonical, valuable URLs rather than every possible parameter combination
- Block low-value, infinite-combination URL patterns (like faceted filters) in robots.txt once you’ve confirmed they add no unique value
- Monitor Search Console’s crawl stats report for sudden drops in crawl requests, which can indicate a crawl budget or server issue
- Cross-check server logs against the sitemap to confirm priority pages are actually being crawled at the frequency you expect
Google’s guidance on crawling and indexing frames sitemaps and robots.txt as complementary discovery tools, not competing ones: sitemaps tell Google what to prioritize, robots.txt tells it what to skip.
Structured data, metadata, and snippet optimization
Metadata and structured data do not directly move rankings, but they control whether your pages earn rich results, better click-through, and growing relevance for AI-driven answer engines.
Start with the basics that are still frequently wrong: unique, descriptive title tags under roughly 60 characters, meta descriptions that summarize the page rather than stuff keywords, and robots meta tags that aren’t accidentally set to noindex on pages you want indexed. A surprising share of indexing problems trace back to a noindex tag left over from a staging environment.
For structured data, prioritize JSON-LD types by what actually appears on the page: Article markup for blog and news content, Product markup for ecommerce listings (price, availability, reviews), FAQ markup for question-and-answer content, and Breadcrumb markup for navigational context. Each type has specific required fields, and missing one can disqualify the page from the corresponding rich result.
A practical validation routine:
- Run every templated page type through the Rich Results Test before launch, not just once in development
- Use a schema validator to catch syntax errors that the Rich Results Test might not flag
- Check Search Console’s Enhancements reports weekly for new structured data errors introduced by content or template changes
- Re-validate after any template update, since a small CSS or component change can silently break JSON-LD output
- Keep metadata and schema in sync, so the title tag, Open Graph tags, and structured data name fields don’t contradict each other
Semantic HTML and clear metadata also matter for how AI retrieval systems parse and cite your content, since those systems rely on the same structural signals as traditional search, according to Search Engine Land’s coverage of technical SEO for AI search. Teams building content specifically for citation in AI answers can find more detail in a guide to optimizing pages for LLM visibility.
Site architecture and internal linking at scale
How your pages link to each other determines how evenly authority and crawl attention get distributed across the site, and faceted navigation is where this most often breaks down.
Keep the hierarchy flat: any important page should be reachable within three or four clicks from the homepage, and URL patterns should stay predictable and consistent within a template type rather than drifting between formats over time.
Follow this order when auditing architecture:
- Crawl the full site with a tool like Screaming Frog to identify orphan pages that have no internal links pointing to them
- Fix broken internal links found during the crawl, prioritizing links from high-authority pages
- Confirm navigation, footer, and sitemap all surface the pages you actually want indexed and ranked
- Audit faceted navigation (filters for size, color, price range) to see which combinations create unique, valuable pages versus which create near-duplicate noise
- Apply canonical tags, parameter handling rules, or robots.txt blocks to faceted URLs based on that value assessment, rather than blocking everything or indexing everything by default
Faceted navigation on large ecommerce catalogs can generate more unique URLs than the rest of the site combined, most of which offer no unique value and simply dilute crawl budget. A practical faceted navigation remediation guide covers the specific canonicalization and parameter-handling patterns that keep filter combinations from overwhelming crawl budget while still letting genuinely valuable filtered pages get indexed.
Security, HTTPS, and server configuration checks
Server-level configuration rarely gets the audit time it deserves, but it affects crawling, indexing, and user trust all at once.
Confirm HTTPS is correctly implemented across the entire site, including subdomains, and that the SSL certificate is valid and not close to expiration. Check for mixed-content warnings, where an HTTPS page still loads an image, script, or stylesheet over plain HTTP, since browsers flag this and it can break parts of the page silently.
Implement HSTS to force HTTPS consistently and eliminate unnecessary redirect hops between HTTP and HTTPS versions of the same URL, since every extra redirect adds latency and a small amount of crawl budget waste. Set cache headers deliberately: static assets like images and fonts should carry long cache lifetimes, while HTML that changes frequently should not be cached aggressively.
A short server-level checklist:
- Confirm every page returns a 200 status for live content and a proper 404 for genuinely missing pages, rather than a “soft 404” that returns 200 with error-page content
- Watch for 500-series errors in both Search Console and server logs, since repeated server errors can cause Google to reduce crawl rate
- Check DNS reliability and response times, since intermittent DNS failures can look identical to a content problem in crawl reports
- Verify redirect rules at the server level don’t create loops between HTTP, HTTPS, www, and non-www versions of the same URL
Audit tools and a step-by-step action plan
Running every check in this article in a useful order matters as much as running it at all. Work through audits in this sequence: indexing status, server errors, performance using field data, mobile usability, JavaScript rendering, metadata, structured data, then internal linking. Each layer depends on the one before it being roughly clean.
- Google Search Console: indexing status, mobile usability, Core Web Vitals report, and manual actions
- Chrome UX Report (CrUX) or a RUM vendor: real-world field data for LCP, INP, and CLS at the 75th percentile
- PageSpeed Insights and Lighthouse: lab diagnostics for root-causing a specific slow page
- WebPageTest: waterfall-level detail on what’s blocking render and when resources load
- Screaming Frog or a similar crawler: site-wide checks for broken links, redirect chains, and orphan pages
- A schema validator and the Rich Results Test: structured data correctness before and after template changes
A simple table of failure thresholds keeps the triage objective rather than subjective:
| Check | Pass threshold | Source |
|---|---|---|
| LCP (field, 75th percentile) | ≤2.5 seconds | Core Web Vitals |
| INP (field, 75th percentile) | ≤200 milliseconds | Optimize INP |
| CLS (field, 75th percentile) | ≤0.1 | Core Web Vitals |
| Server error rate | No recurring 5xx spikes in logs | Crawling and indexing |
Once the audit surfaces issues, prioritize fixes by the overlap of traffic impact and fix difficulty: a quick canonical tag correction on a high-traffic template outranks a complex rendering migration on a low-traffic section. After each fix ships, return to field data rather than a single lab test to confirm the fix actually changed real user experience, then set a recurring review (monthly for Core Web Vitals, weekly for Search Console errors during active remediation) so regressions get caught before they compound. This staged approach, fixing crawlability and indexation before performance and before structured data enhancements, mirrors the layered priority order described in Search Engine Land’s technical SEO hierarchy of needs.
Practical engineering notes on remediation trade-offs
Technical SEO fixes split into two categories: changes a marketer can make directly in a CMS, and changes that require engineering time because they touch rendering, infrastructure, or build pipelines. Knowing which bucket a fix falls into before committing to a timeline avoids a lot of friction between SEO and engineering teams.
Server-side rendering versus client-side rendering is the clearest example of this trade-off. Client-side rendering is often faster to ship for a development team, but it pushes rendering work into the browser and into Google’s rendering queue, adding delay before new content gets indexed. Server-side or hybrid rendering costs more engineering time upfront but removes that delay and gives direct control over what crawlers receive.
Platform-level fixes tend to pay off longer than page-level ones:
- A content delivery network reduces server response time globally, which helps LCP on every page at once
- Long-lived cache headers on static assets reduce repeat-visit load time without any page-level work
- Content fingerprinting (versioned filenames for JS and CSS) lets you cache aggressively while still forcing updates when files actually change
Pro Tip: Before scoping a rendering migration, measure how much indexing delay is actually costing you in the field; sometimes a caching and CDN fix closes most of the gap at a fraction of the engineering cost.
Handling duplicate content issues beyond canonical tags
Canonical tags solve the signal problem, telling Google which version you prefer, but they don’t solve the root cause of why duplicates exist in the first place. A site that keeps generating new duplicate URLs will keep needing new canonical tags indefinitely.
Common sources of duplication beyond simple parameter variants include printer-friendly page versions, session IDs appended to URLs, staging or testing subdomains that got crawled before being blocked, and content syndicated to other domains without a cross-domain canonical pointing back to the original. Each needs a different fix: eliminate session IDs from URLs entirely rather than canonicalizing around them, noindex or remove printer-friendly duplicates, and make sure any syndication partner includes a canonical link back to your original article.
Google’s guidance on consolidating duplicate URLs recommends stacking signals consistently: matching redirects, rel=canonical tags, sitemap entries, and internal links all pointing the same direction, since contradictory signals across those methods make it harder for Google to pick the version you actually want indexed. The fix that lasts is removing the source of duplication, not just labeling it after the fact.
Pagination and hreflang for international and multi-language sites
Paginated series (page 2, page 3 of a category or blog archive) and multi-language sites both create a different kind of duplication risk, one based on sequence and locale rather than simple copy-paste content.
For pagination, each page in a series should have its own self-referential canonical tag rather than pointing back to page one, since each page typically contains distinct products or posts. Avoid “view-all” canonicalization unless the view-all page actually loads acceptably fast, since forcing a canonical to a slow page can hurt the metric you’re trying to protect.
For multi-language and multi-region sites, hreflang tags tell search engines which language or regional version of a page to show to which audience. Every page in a hreflang set needs to reference every other version, including itself, and the return tags need to match exactly. A one-directional hreflang reference (page A points to page B, but B doesn’t point back to A) is a common error that makes the entire set unreliable. Keep hreflang values to valid language and region codes, and don’t assume a single sitemap can silently handle this. Dedicated hreflang sitemap entries or in-page tags, implemented consistently across the entire set, are what actually make the signal usable.
How we run a technical SEO audit: a practitioner’s note
Our own view, shaped by running these audits alongside engineering teams, is that the order matters more than the tool list. Start with field data, not lab scores, because lab tools measure a single simulated session while field data reflects what your actual visitors experienced. Indexing and crawlability come before performance, and performance comes before structured data polish, because a perfectly marked-up page that isn’t indexed helps nobody.
Realistic remediation timelines vary by fix type: a canonical or metadata correction can ship in days, while a rendering migration or CDN rollout often takes weeks of coordinated engineering work. The clearest sign an audit worked is field data moving in Search Console and CrUX over the following few monitoring cycles, not a single lab score improving overnight.
— Service
Running this checklist tells you what’s broken. Fixing a rendering pipeline, a CDN misconfiguration, or a faceted navigation system at scale is a different kind of work, and it’s a type of work that marketing and engineering teams often need to handle together.

We embed directly with your team rather than handing you a report and disappearing, which means the people who diagnosed the rendering delay or the redirect chain are the same people who fix it. That matters most for:
- Multi-location brands managing the same template across many regional pages
- B2B SaaS teams untangling a JavaScript framework migration that stalled indexing
- Enterprise commerce organizations with faceted navigation generating more URLs than their crawl budget can handle
Engagements run through Core, Growth, and Scale capacity plans, structured around the priorities on your roadmap rather than a fixed project scope. If remediation work has stalled on your team’s backlog, that’s a reasonable place to start the conversation.
FAQ
What does technical SEO consist of?
Technical SEO covers everything that affects whether search engines can crawl, render, index, and understand your pages: crawlability, site speed and Core Web Vitals, mobile usability, JavaScript rendering, URL structure, canonicalization, structured data, and server configuration. It sits apart from content and link-building work, though all three layers affect rankings together.
How to do a technical SEO audit?
Work in a fixed order: confirm indexing status and server error rates first, then check Core Web Vitals using field data such as CrUX, then review mobile usability, JavaScript rendering, metadata, and structured data. Fix issues in that same order, since problems in earlier layers can mask or distort the results of fixes made later.
What is the difference between SEO and technical SEO?
SEO is the full discipline of improving organic search visibility, including content quality, keyword targeting, and link building. Technical SEO is the subset focused specifically on crawlability, indexing, rendering, and site infrastructure, the mechanical layer that determines whether your content can be found and understood at all.
How often should a technical SEO audit be run?
A full audit covering every section of this checklist is reasonable once or twice a year, but Core Web Vitals and Search Console error reports are worth checking monthly, since performance regressions and indexing errors tend to surface gradually between full audits.
Should small sites follow the same checklist as large enterprise sites?
The same checks apply, but priority shifts: a small site rarely has crawl budget problems or faceted navigation issues, so indexing, Core Web Vitals, and mobile usability deliver most of the value, while sitemap splitting and crawl budget management matter far more as catalog size grows.