Run a bounded technical crawl
Crawl a capped same-site sample and review raw-HTML findings without implying full-site coverage.
Free, no signup. Crawl up to 150 exact-origin HTML pages, then check response status and headers for up to 100 exact-origin assets. Group raw-HTML SEO findings, inspect affected URLs, and export Pages, Links, Redirects, Assets, or Issues CSVs.
Configurations and up to ten previous run summaries stay in this browser. Nothing runs on a schedule or in the background. Do not put secrets in URLs.
Checks run from our server; we fetch the URL you enter and don't keep the results. Up to 150 public HTML pages and status/header evidence for up to 100 exact-origin assets are retrieved through bounded, rate-limited proxies. Asset bodies are cancelled and never retained. Results stay in this browser unless you export them. Anonymous run-level outcome counters may be used for aggregate research; URLs, domains, IPs, and identifiers are never included, and no statistic is released below 100 runs.
82/100 raw-HTML site score
Crawled 24 of 31 discovered URLs (partial — page cap or policy exclusions would be named here).
title/missing · warn · 2 affected · checked on 24/24
Not evaluated: rendered JavaScript and AI checks.
This is illustrative output, not an audit of a real site and not a search-engine score.
+ saves the current site or page. Use ☆ beside any saved site, page, or list to favorite it. Recent check history appears below.
Target filled from your local choices.
Saved targets, named lists, and recent check summaries remain only in this browser.
Affected share uses detector-page pairs where each check had enough evidence. Crawl depth uses the bounded pages observed in this run.
Cells show observed affected URLs inside each inferred URL cohort. A zero means no finding was observed in the bounded cohort, not that every possible check passed.
These findings depend on relationships across pages, crawl inventory, or asset responses, so they appear before the page-detector ledger.
Problems is the default view. All checks adds passes, notices, and checks that lacked enough evidence.
Scout only scores evidence-backed problems whose likely intent and consequence can be established from this bounded raw-HTML crawl. Configuration choices, informational opportunities, uncertain intent, and missing evidence stay visible without becoming penalties. “Not evaluated” never means “passed.”
Outside this tier: rendered JavaScript, authenticated pages, and AI checks are intentionally not run. Smart JS and Full Render are the rendered tiers.
A protected queue crawls exact-origin HTTPS pages from a root or bounded sitemap inventory, checks robots policy, normalizes discovered links, and caps trap patterns. Each usable raw response becomes a page record for deterministic detectors. The rollup preserves evaluation coverage and creates exportable issue evidence.
The crawl is bounded, raw-HTML only, and exact-origin; it cannot prove full site coverage or orphan status from a root crawl. Asset checks do not download or analyze file bodies, calculate transferred bytes, or evaluate third-party resources. Scout does not render JavaScript, honor crawl-delay, authenticate, run AI checks, call GSC, or emulate a search crawler’s index. Security findings are conservative review prompts, not proof that a site is exploitable.
It fetches up to 150 exact-origin HTML pages, then runs status-only checks for up to 100 exact-origin HTTPS assetsHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.' extracted from complete raw HTML. Page and asset coverage are reported separately.
No. Findings use fetched raw HTML. JavaScript-rendered content and links are explicitly not evaluated.
The page crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. applies Allow and Disallow rules for its crawl policy. This tool does not honor crawl-delayCrawl-delay is a non-standard robots.txt directive that asks bots to wait between successive page fetches to ease server load. Google has ignored it since September 1, 2019 and Yandex dropped it on February 22, 2018; Bing and many other crawlers still honor it, but Bing interprets the value differently than the plain reading., which it states in the scope note.
It is a technical condition score derived from the impact and affected scale of evaluated, score-eligible page issues in the bounded sample. Asset findings are reported separately. It is not a Google score or a complete measure of site quality.
Yes. Export a bounded local JSON snapshot and import it after another run. Coverage shifts can affect the comparison.
Upvote what you want most. New ideas can be submitted from the floating Feedback menu; requests appear here once approved, and the most-wanted rise to the top.
You won't be emailed about that request anymore.
Loading…
New requests are reviewed before they appear here.
Where this tool helps
Crawl a capped same-site sample and review raw-HTML findings without implying full-site coverage.
Group recurring status, indexability, metadata, link, and asset issues by detector and affected URL set.
Use snapshots to see which detector and inferred-cohort findings appeared, disappeared, or changed.
Export CSV and finding packets with provenance so engineers can reproduce and verify the prioritized work.
Watch the full workflow
Scout Site Audit Free runs a bounded exact-origin crawl, groups deterministic page, cross-page, and asset findings, exposes coverage, prioritizes evidence-backed issues, and exports auditable reports. I’ll show root and sitemap setup, crawl controls, a clearly fictional dashboard, score and coverage interpretation, exports and comparisons, limitations, and the best next steps.
Use Scout for launch checks, migrations, template reviews, recurring spot audits, and evidence handoffs. The free tier fetches up to one hundred fifty exact-origin H-T-M-L pages and status-checks up to one hundred eligible assets.
Enter an H-T-T-P-S root for link discovery, or identify an X-M-L sitemap as the bounded starting inventory. The crawl stays on the exact origin and applies matching robots Allow and Disallow rules.
Optional same-origin manual seeds cover known important pages. Watchlists and up to ten previous run summaries stay in this browser; they do not schedule background crawls. Never put secrets in U-R-Ls.
This walkthrough fills fictional example dot com but does not start a crawl or contact a site. A real run requires the anti-abuse check, shows live progress, supports abort, and separates the page and asset phases.
This illustrative summary combines an evidence-backed raw-H-T-M-L condition score with explicit page and asset coverage. It is not a Google score, and partial coverage narrows every clean, absence, and site-wide claim.
Category affected-share uses detector-page pairs with enough evidence. Crawl depth describes only observed bounded pages. Use these charts to locate patterns, then inspect the underlying U-R-L evidence.
Scout ranks errors before warnings, then uses disclosed impact and affected scale. Confirm page intent and shared causes before changing templates; the list is an investigation order, not an automatic fix queue.
Cross-page findings depend on crawl inventory and link relationships. Asset findings use response status and allowlisted headers only; bodies are cancelled and third-party resources are not fetched. Page checks remain grouped separately.
Problems view focuses on actionable warnings and errors; All checks adds passes, notices, and unavailable evidence. Expand rows to inspect affected U-R-Ls and detector coverage. Not evaluated never means passed.
Export Pages, Links, Redirects, Assets, or Issues C-S-Vs, human-readable findings, a printable report, a privacy-sanitized snapshot, or a graph handoff. Import a prior snapshot only after comparing coverage shifts.
The crawl cannot prove full coverage or orphan status, render JavaScript, authenticate, honor crawl-delay, run A-I checks, call Search Console, or emulate a search index. Asset checks do not analyze bodies or third-party resources.
Confirm the top findings on representative pages, fix shared root causes, export the report, and rerun the same seed and scope. Compare both coverage and evidence, then escalate JavaScript, authenticated, or deeper page questions to the appropriate tool.
Start with the prioritized evidence, confirm intent on representative affected pages, and fix the shared template or engine where appropriate. Export the bounded report, recrawl the same scope, and compare coverage before claiming improvement. Use Page Report for deeper single-U-R-L work and rendered or authenticated tools for evidence this raw-H-T-M-L tier cannot observe.