Make AI results inspectable
Agent evaluation, workflow design and exact search, with explicit tasks, replayable records and independent verification.
A plausible result needs a way to be checked. For agents, that means retaining the task, tool actions, observations and failure conditions. For an exact search, it means distinguishing a candidate from an independently verified solution.
Start with the workflow question in The first AI managers. Follow it into the ViralBench harness and replayable traces, then compare those checks with a solver whose output has a precise verification condition.
Question paths added .
The Shopkeeper in the Machine
Examine how people assign, supervise and evaluate AI work inside a workflow.
14 min readBeyond the Leaderboard
Inspect the harness, task records and evaluation boundaries behind an agent experiment.
23 min readReplayable traces for AI agents.
Turn an agent result into a record that another evaluator can inspect and replay.
6 min readVerifying a 54-move exact search.
Follow an exact search from constraints and candidates to a separately checked answer.
8 min read