Opening the page…

Sulayman Bowles

Make AI results inspectable

Agent evaluation, workflow design and exact search, with explicit tasks, replayable records and independent verification.

A plausible result needs a way to be checked. For agents, that means retaining the task, tool actions, observations and failure conditions. For an exact search, it means distinguishing a candidate from an independently verified solution.

Start with the workflow question in The first AI managers. Follow it into the ViralBench harness and replayable traces, then compare those checks with a solver whose output has a precise verification condition.

Question paths added .

  1. The Shopkeeper in the Machine

    Examine how people assign, supervise and evaluate AI work inside a workflow.

    14 min read
  2. Beyond the Leaderboard

    Inspect the harness, task records and evaluation boundaries behind an agent experiment.

    23 min read
  3. Replayable traces for AI agents.

    Turn an agent result into a record that another evaluator can inspect and replay.

    6 min read
  4. Verifying a 54-move exact search.

    Follow an exact search from constraints and candidates to a separately checked answer.

    8 min read

Browse all writing