All insights

Who reviews the agent that reviewed the agent?

Another review agent can be part of the design. Trust in the work requires knowing what was checked, against which evidence, and how a finding becomes an improvement.

Illustration: a professional reviewing a report, with a close-up of an evidence checklist.

A quality-review agent sounds like a natural role in a digital team. Its title, however, does not define a quality threshold. Without clear failure criteria and evidence of acceptance, even a long chain of reviewers can leave the actual decision unclear.

Define acceptable work in the language of the task

We suggest describing quality in terms of the assignment. A document-based draft may need a source for each material claim and explicit missing-information markers. A campaign package may need consistency with the brief and support for its promises. A financial workflow may need period and source alignment alongside calculation checks.

This definition helps separate structured checks from questions requiring interpretation. A sum, required field, or working reference may support a defined check. The clarity of an explanation or meaning of a contradiction may require judgment. Each kind of review should use an appropriate mechanism.

A finding should change the state of work

Suppose a reviewer finds a claim based on an outdated version. Its output should explain the finding, identify the source, and return the work to a role capable of correcting it. The process also needs to define a sufficient correction and when to escalate to a person.

In AOS, agent managers hold the execution-and-review loop within the assignment. A finding connects to an output, a revision, and a decision. The number of agents is not a useful quality measure by itself. We want to understand why work was accepted and what remains unresolved.

Continuous improvement is a change that can be evaluated

Work findings can suggest a refinement to context, memory, tools, or skills. Examples include checking a document’s validity, clarifying a tool description, or specifying when an additional source is needed. We treat the proposal as a change requiring evaluation, a recorded history, and permission boundaries.

In practical terms, learning does not automatically expand authority. An agent may propose a more precise instruction, while organizational and system policy determine whether and where it applies. A change that helps one case should also be examined in situations where it might interfere.

Anthropic’s guidance on agent tools emphasizes realistic evaluation tasks with verifiable outcomes, and testing changes on cases held apart from the improvement process. Our takeaway is to evaluate a proposed change beyond the example that originally motivated it. Writing tools for agents — Anthropic

Human approval needs a decision package

An approval control is useful when the person understands what they are approving. We suggest presenting the proposed action, output, sources, limitations, and unresolved questions together. The person should be able to return the work with a focused request when needed.

Approval also needs a clear scope. Permission to prepare a draft differs from permission to send it. Approval of a change package differs from approval to release it into production. That distinction belongs inside the workflow, so the next agent receives authority it can interpret.

A practical management question

When someone demonstrates an agent team, ask to follow one error. Who discovered it? Who became responsible for the correction? Which evidence showed the correction was sufficient? What changed to reduce recurrence? Concrete answers explain the review mechanism better than a broad statement about oversight.

For GOBOOST, this is where technology and management connect: a system organizing work, alongside a process defining quality and authority. It creates a basis for expanding the use of agents while understanding what they do, how their work is checked, and how the next assignment can improve.

TAKE IT TO WORK

Follow one error through correction and evaluated improvement. It is a practical way to examine how the agent team is managed.

Explore AI teams in action