AI agents increasingly work in teams, and the human role shifts from producing work to checking it. This shift rests on two assumptions: that checking costs less than doing, and that a person who checks can tell correct work from work that only looks correct. Neither has been measured systematically. This thesis makes verification itself the object of study - the process by which a human or automated verifier establishes whether an agent team's output meets the task requirements and constraints.
The first line of work measures the costs of verification. Controlled experiments with human participants will vary the number of agents, the length of the trajectory, how much of the agents' communication happens in parallel, and how tightly their decisions depend on one another. The experiments separate three failure types: local errors, errors that propagate, and team-level failures in which every agent behaves plausibly but the joint outcome is wrong.
The second line asks where verification stops being cheaper than production. It tests whether the cost of checking a trajectory, and the chance of catching an error in it, can be predicted from properties of the trajectory itself. It also develops a formal model that treats oversight as the allocation of a limited attention budget across agents whose errors are correlated, and mechanisms that adapt how agents communicate, how they explain themselves, and how much they act on their own.
The third line tests the reverse pressure: whether agents optimised against an oversight signal learn to produce work that is deliberately hard to check, and which oversight signals survive that pressure.
These research lines are intended to establish verification as a measurable object rather than an assumption: an empirical account of what checking agent work actually costs, a formal model of oversight under a bounded attention budget, an estimator that makes verification cost predictable in advance, mechanisms that adapt agent autonomy to the person doing the checking, and a benchmark that scores the verifier rather than the agent. Domains will be selected to span work that differs in how directly it can be checked, from tasks with a verifiable ground truth to tasks where correctness is a matter of judgement.