Opening the page…

Sulayman Bowles

Build search systems you can verify

Robots rules, crawl queues, rendered pages, canonical identity and release checks, connected through evidence rather than dashboard labels.

A useful crawl records what was requested, what came back and which interpretation produced a finding. It also states what those observations cannot establish. A user-agent string does not prove identity; a robots rule does not protect private content; a canonical declaration does not prove Google selected it.

Begin with the control you need, then follow the implementation through capture, storage, interpretation and release. Atlas provides the worked system; the individual essays explain its engineering decisions and their limits.

Question paths added .

  1. Robots.txt is not access control.

    Choose the right control for crawling, indexing, private access and expensive requests.

    6 min read
  2. AI crawlers: search, training, and access.

    Distinguish crawler purposes and claimed identities before deciding on a policy.

    4 min read
  3. Building Atlas

    See how capture records, derived findings and a reader-facing interface fit together in a working project.

    4 min read
  4. The crawl frontier is a state machine.

    Design scheduling, retries and terminal states so a crawl can resume without losing its history.

    6 min read
  5. SQLite for crawl pipelines.

    Keep queue state and observations in a storage model that supports inspection and replay.

    6 min read
  6. Raw HTML and rendered DOM are separate evidence.

    Compare a response body with the rendered document while keeping each observation attributable.

    6 min read
  7. Canonicalization is a graph problem.

    Resolve declarations and redirects as a graph instead of checking isolated tags.

    6 min read
  8. Internal links are a retrieval graph.

    Use internal links to examine reachable documents and retrieval paths.

    6 min read
  9. Structured data without content drift.

    Keep machine-readable claims aligned with the content visitors can actually read.

    5 min read
  10. Audit findings are derived records.

    Make findings reproducible from observations, rule versions and explicit evidence.

    6 min read
  11. Executable gates for SEO migrations.

    Check routes, redirects and discovery before a migration becomes a search regression.

    6 min read
  12. Technical SEO as public data infrastructure.

    Understand the boundary between public technical observations and private Search Console measurements.

    5 min read
  13. Canonical identity beats more content.

    Keep a person’s identity consistent across URLs, metadata and visible biography.

    5 min read

Browse all writing