Build search systems you can verify
Robots rules, crawl queues, rendered pages, canonical identity and release checks, connected through evidence rather than dashboard labels.
A useful crawl records what was requested, what came back and which interpretation produced a finding. It also states what those observations cannot establish. A user-agent string does not prove identity; a robots rule does not protect private content; a canonical declaration does not prove Google selected it.
Begin with the control you need, then follow the implementation through capture, storage, interpretation and release. Atlas provides the worked system; the individual essays explain its engineering decisions and their limits.
Question paths added .
Robots.txt is not access control.
Choose the right control for crawling, indexing, private access and expensive requests.
6 min readAI crawlers: search, training, and access.
Distinguish crawler purposes and claimed identities before deciding on a policy.
4 min readBuilding Atlas
See how capture records, derived findings and a reader-facing interface fit together in a working project.
4 min readThe crawl frontier is a state machine.
Design scheduling, retries and terminal states so a crawl can resume without losing its history.
6 min readSQLite for crawl pipelines.
Keep queue state and observations in a storage model that supports inspection and replay.
6 min readRaw HTML and rendered DOM are separate evidence.
Compare a response body with the rendered document while keeping each observation attributable.
6 min readCanonicalization is a graph problem.
Resolve declarations and redirects as a graph instead of checking isolated tags.
6 min readInternal links are a retrieval graph.
Use internal links to examine reachable documents and retrieval paths.
6 min readStructured data without content drift.
Keep machine-readable claims aligned with the content visitors can actually read.
5 min readAudit findings are derived records.
Make findings reproducible from observations, rule versions and explicit evidence.
6 min readExecutable gates for SEO migrations.
Check routes, redirects and discovery before a migration becomes a search regression.
6 min readTechnical SEO as public data infrastructure.
Understand the boundary between public technical observations and private Search Console measurements.
5 min readCanonical identity beats more content.
Keep a person’s identity consistent across URLs, metadata and visible biography.
5 min read