SULAYMAN BOWLES Technical SEO · AI Systems · Finance Research INDEX  +

Research / public archive

One archive. Clear categories.

Technical SEO research and AI systems notes from Sulayman Bowles on crawlability, crawler policy, Atlas, public data, identity, markets, and evidence.

Opening the current route
SULAYMAN BOWLES Technical SEO · AI Systems · Finance Research

Research Notes

Selected notes on search systems, crawlability, Atlas, public data, and markets work.

Four Research Categories

Crawler engineering

Crawl frontiers, robots policy, rendering evidence, and retrieval systems.

Technical SEO

Canonicals, internal links, structured data, migrations, and identity consistency.

Data & AI systems

Derived records, SQLite pipelines, agent evaluation, and AI operations.

Finance & ownership

Capital stacks, infrastructure, market mechanics, and local ownership.

Public studies & archive

Atlas evidence, the Austin pilot, and clearly labeled retired methodology.

25 Notes and Artifacts

The Crawl Frontier Is a State Machine, Not a Queue

A practical architecture for URL identity, admission control, host politeness, bounded retries, crawl traps, and evidence-preserving frontier transitions.

Raw HTML and Rendered DOM Are Separate Evidence

How to capture, compare, and qualify transport source, browser output, dependent requests, runtime failures, and render completeness without treating screenshots as data.

Canonicalization Is a Graph Consistency Problem

A systems method for finding conflicting canonicals, redirect chains, cycles, duplicate clusters, and sitemap disagreements before they become indexation ambiguity.

Internal Links Are a Directed Retrieval Graph

A graph-first method for crawlable links, reachability, depth, contextual edges, orphan diagnosis, and repair plans that preserve user intent.

Robots.txt Is a Courtesy Layer, Not Access Control

A precise model for separating crawler requests, indexing directives, authentication, authorization, rate controls, and evidence of enforcement.

Structured Data Without Content Drift

A typed-content architecture for keeping visible pages, metadata, JSON-LD, sitemaps, and exports consistent through generation, invariants, and production tests.

Audit Findings Should Be Derived Records

A data model that separates captured observations, artifacts, rule evaluations, findings, confidence, review state, and recommendations without losing lineage.

Replayable Traces for Evaluating Tool-Using AI Agents

An evaluation architecture for tasks, trials, observable trajectories, controlled environments, layered graders, repeated runs, and evidence-gated promotion.

SQLite for Crawl Pipelines: Idempotency, WAL, and Bounded Concurrency

A storage architecture for URL identity, append-only attempts, transactional batches, upserts, one-writer discipline, WAL checkpoints, integrity checks, and portable exports.

Technical SEO Migrations Need Executable Release Gates

A fail-closed migration method for URL manifests, redirect graphs, canonical output, internal links, sitemaps, rendered content, launch sequencing, and post-release evidence.

The First AI Managers

A 30-case review of AI-operated businesses that separates live operations, pilots, simulations, and vendor claims.

Beyond the Leaderboard: ViralBench + Codex

A code-level design for traces, replay, controlled trials, and a bounded engineering loop around a live marketing agent.

AI Crawler Robots.txt Guide: GPTBot, OAI-SearchBot, ClaudeBot and PerplexityBot

Copy-ready host rules for eight named product tokens, with exact Allow, Disallow, sitemap, and release checks.

Technical SEO as Public Data Infrastructure

A systems essay on how URLs become crawlable, renderable, attributable, and exportable public records.

Canonical Identity Beats More Content

An operational playbook for reconciling domains, profile pages, resumes, sameAs links, and external biographies.

Who Owns Austin’s Home-Service Companies?

A verified map of the parent companies, private-equity sponsors, public corporations, franchises, and local operators behind 67 home-service brands advertising across Austin.

What Happens When an Index Decides a Company Matters?

How private index rules become public market orders—and why inclusion can move ownership, liquidity, and price without guaranteeing permanent value.

How Airlines Borrow Against Loyalty Programs

How airline loyalty program financing turns bank payments, co-brand contracts, controlled accounts, and loyalty IP into collateral—without treating points as deposits.

Where Do Online Returns Go? Inside Reverse Logistics

What happens to online returns after a refund: a product-level model of restocking, open-box resale, refurbishment, liquidation, recycling, destruction, and returnless refunds.

Hardware Startup Financing: Five Capital Stacks

How hardware startup financing combines venture equity, equipment finance, asset-backed debt, leases, customer capital, public support, and project finance across five companies.

Who Owns West Campus Student Housing?

The captive economics of student housing around the University of Texas at Austin, traced from one student installment through property operations, capital structure, and first loss.

Who Funds Waymo’s Hardware?

A Waymo case study of who finances the vehicles and infrastructure, who owns the risk, and who absorbs the loss when utilization or technology fails.

Who Owns the Toll Roads in Texas? Ownership, Operators, and Economics

The state usually owns the pavement. Contracts decide who controls toll revenue, who gets paid first, and who absorbs the loss.

Atlas Open Corpus Demonstration

A versioned public-corpus run showing source and rendered states, traceable findings, confidence, and export excerpts.

Austin Crawlability Pilot

A bounded 12-site public-homepage pilot with a dated cutoff, public CSV, and explicit measurement gaps.

From Research to Implementation