Gauntlet Loop: AI without criteria just repeats the mistake faster

Understand how to benefit from the prompt technique that elevates the level of LLMs.
Understand the Gauntlet Loop and how to use AI agents with criteria, independent evaluation, cost limits, and connection to real-world outcomes.

The team produces campaigns in hours.

Build a landing page in the morning, fifteen ads before lunch, and a market analysis at the end of the day.

The volume increased. The speed did too.

But the review keeps taking forever. The message changes from one piece to another. The form captures data that sales doesn't use. The campaign promises one thing and the page explains another. When the material reaches leadership, a sequence of adjustments that no one had anticipated begins.

AI delivered more.

This does not mean that they delivered better.

The bottleneck is shifting. For years, the problem was producing. Now, in many operations, the problem is defining what deserves to be published.

In the article about Vibe Marketing, I discussed the new way of producing with AI, with faster cycles and a different relationship between idea and execution. Here, the discussion is different.

If producing became cheap, who protects the standard?

The next AI leap in marketing does not come from a more ingenious prompt. It comes from a system that separates production, judgment, and evidence. And that knows where the evaluation of the artifact ends and the market response begins.

This is where the Gauntlet Loop deserves attention.

Not as a magic formula.

As a management provocation.

What is the Gauntlet Loop

"Gauntlet Loop" is the name Matt Shumer began using to describe the process employed in the experiment Claude of Duty, a first-person shooter created using an AI agent framework.

According to Shumer's account, a primary agent received a goal and a concrete quality reference. Then, it broke the work down into smaller parts, distributed the construction among agents, and placed each part before independent critics.

The process followed this logic:

  1. Set a goal.
  2. Present a concrete reference of what quality means.
  3. Divide the work into parts that can be evaluated separately.
  4. Assign each part to a building agent.
  5. Use another agent, with separate context, to evaluate the result.
  6. Compare the artifact with the reference.
  7. Identify the biggest difference.
  8. Correct and re-evaluate.

Shumer summarizes the core of this dynamic in four steps: 

  • divide, 
  • build, 
  • to judge 
  • and repeat.

The key point is not just repetition. It is about preventing the same person who produced the work from being the sole authority on the quality of that work.

Gauntlet Loop. More autonomy requires more governance.

This didn't come out of nowhere. 

Anthropic had already described a similar pattern in Building Effective AI Agents, under the name of evaluator-optimizer. 

One person creates, another evaluates, and the cycle continues as long as the review leads to measurable improvement.

"Gauntlet Loop" is a recent term for a specific combination of these ideas, expanded to include task decomposition, external references, and multiple agents. 

It is not a scientifically consolidated methodology. 

There is no basis to treat it as a guarantee of results.

The experiment itself helps keep the hype in check.

No evaluation report published in the repository, the goal was to match the visual quality of a modern Call of Duty game. That didn't happen.

The critics' ratings were as follows: 3,59, 4,14, 4,05 e 5,05 in a scale of ten

Two screenshots reached the “close” rating. 

The others remained classified as “amateur.” In the blind comparisons, the critics chose the original game's images in every round.

There was partial improvement.

There was no victory over the reference.

In the published report, the artifact improved over the rounds. 

The case suggests that comparison and criticism can help, but it does not isolate the cause of the improvement nor prove overall effectiveness.

Nor does it demonstrate that autonomous agents can achieve any standard, in any context, simply because they continue working for a longer period of time.

What really changes for leadership?

The superficial reading of the Gauntlet Loop is that AI has become more autonomous.

The useful interpretation is that autonomy requires more governance.

The human can stop reviewing every detail of every version. 

But it remains responsible for decisions the agent should not make alone:

  • What is the real goal.
  • Which reference represents quality.
  • Which criteria cannot be violated.
  • How much time and money the process can consume.
  • When the improvement stopped justifying another round.
  • Who approves the publication.
  • What happens when the raters disagree.

Autonomy without reference produces volume.

Unbounded reference can produce an infinite bill.

In the original text about the Gauntlet Loop, Shumer argues that the cycle should continue until the result reaches the standard or until a person decides to stop it. This may make sense in an experiment. 

Inside a company, it is insufficient.

My adaptation for a real operation adds controls that should not be optional.

Time and cost cap

Every cycle needs a budget, a deadline, and a consumption limit. The fifth round can improve the artifact. It can also cost more than that improvement is worth.

Recurring column

The criteria need to exist before the evaluation. If the rubric changes every time the agent encounters a difficulty, the system is not improving. It is moving the finish line.

Leadership may recalibrate the rubric between rounds, provided it records the change and restarts the comparison under the new criterion. 

The agent must not alter it alone to facilitate its own approval.

Maximum iterations

Even with a financial cap, it is worth establishing an initial number of rounds. If the result does not converge, the task goes back to human decision.

Change log

Each round needs to show what changed, why it changed, and what criterion justified the alteration. Without history, you have no learning. You only have versions.

Invariants

Brand, compliance, legal requirements, security, value proposition, and business rules must be treated as constraints. An agent cannot sacrifice a mandatory requirement in order to score points in aesthetics or projected conversion.

Safe return

If a round worsens the result, the team needs to roll back to the last approved version. Evolution without rollback capability is operational risk.

Human approval

AI can recommend that the asset is ready. 

The decision to publish, launch, invest budget, or change the objective remains with a responsible person.

The role of leadership does not disappear. It levels up.

Exit phrase correction and enter the system definition that decides what is acceptable.

B2B Governance for AI Use

Artifact loop is not a business loop

Here is the most important separation.

An agent can evaluate a landing page before publication. 

You can check clarity, consistency, brand adherence, form operation and presence of evidence.

But he cannot, just by looking at the page, conclude that it will generate a pipeline.

This only appears after real buyers interact with the asset and the commercial process records what happened.

Dimension

Artifact loop

Business loop

Evaluated object

Page, ad, code, report, or analysis

Conversion, MQL to SQL, pipeline, cycle, and revenue

Signal speed

Quick, before or right after production

Slow, following exposure to the market

Source of comparison

Heading, reference, tests, and requirements

CRM, Buyer Behavior, and Business Results

Role of AI

Build, inspect, compare, and suggest corrections

Support analysis and pattern identification

Person with final responsibility

Owner of the artifact

Marketing, Sales, and Revenue Leadership

Stop criterion

Minimum standard met within the limits

Sufficient evidence to continue, adjust, or discontinue

The complete flow looks like this:

quality before publication → market testing → operational results → lessons learned incorporated into the standard

The first part reduces avoidable errors.

The second one shows whether the asset works in the actual system.

Gauntlet Loop: AI without criteria just repeats the mistake faster

Confusing the two creates a new version of the old problem. 

  • The team optimizes what it can measure quickly and calls that a result.
  • An ad can receive a high rating from an evaluator and attract the wrong audience.
  • A landing page can fulfill the entire brief and still generate leads that do not advance.
  • A report can be flawless and support a bad decision because it was based on incomplete data.

The Gauntlet Loop does not generate revenue.
Degasperi, Israel 

It can improve the reliability of the assets that influence the revenue system. The difference may seem small in the sentence, but it is enormous in management.

A B2B example: the campaign landing page

Imagine a B2B company launching a campaign targeting operations decision-makers.

The goal shouldn't just be to “create a high-conversion landing page.” That’s too general and leaves room for optimizing a single metric.

A better goal would be:

“Create a landing page for operations directors at companies with more than 200 employees that clearly explains the problem, presents verifiable evidence, collects the necessary data for qualification, and correctly transfers the lead to the CRM.”

This work can be divided into four sections:

  • Main promise.
  • Evidence and argumentation.
  • Form and Qualifications.
  • Data transfer to the CRM.

The builder agent creates the first version.

The critic receives the page, the reference is a rubric. 

He evaluates:

  • Is the message aligned with the ICP?
  • Is the promise specific, or could it belong to any company?
  • Do the proofs support the promise?
  • Does the hierarchy allow understanding the offer quickly?
  • Does the form capture the fields used by sales?
  • Were consent, privacy, and brand requirements respected?
  • Do source, campaign, and other information arrive correctly in the CRM?
  • Is there any contradiction between the ad, the landing page, and the sales approach?

If the asset fails, it goes back for correction.

If it meets the standard, it proceeds to human testing.

The AI judge ends here.

Before publishing, the team records the internal approval rate and the amount of rework.

After publication, another cycle begins. 

Now, the signs are:

  • Page conversion.
  • MQL to SQL.
  • Cost per opportunity.
  • Pipeline affected.
  • Revenue, respecting the sales cycle.

This reading needs to be discussed with CRM and RevOps

If marketing celebrates form conversions while sales discards almost all leads, the artifact failed the wrong test.

It's the same problem that occurs when the acquisition does not become revenue

There is no shortage of leads. 

There is a lack of connection between leads, qualification, CRM, the sales pipeline, and results.

When Does the Gauntlet Lopp Make Sense for Companies?

When the Gauntlet Loop makes sense

The method tends to be useful when the work has four characteristics:

  • The deliverable is a physical item and can be inspected.
  • The objective can be divided without destroying the coherence of the whole.
  • There is a reliable reference or rubric.
  • The cost of the error warrants further evaluation.

Pages, campaigns, reports, analyses, automation flows, and commercial materials can benefit from this model.

More Not every task needs an agent system.

A simple change may cost less if it is made and reviewed directly by one person.

Anthropic itself recommends starting with the simplest solution and accepting greater complexity only when the expected benefit justifies the cost and latency.

I would also avoid the Gauntlet Loop when:

  • The decision is irreversible.
  • There is no reliable source.
  • The criteria are predominantly political or subjective.
  • The agent would have access to publish or modify critical systems without approval.
  • The coordination cost exceeds the value of the asset.

There are also less obvious limits.

An evaluator may carry the same biases as the developer. 

A bad rubric can turn compliance into false quality. 

Improved parts separately can form an inconsistent whole. And when a metric becomes a target, the agent can learn to improve the score without improving what actually matters.

The Claude of Duty repository shows a relevant example. 

According to the published analysis, parallel rounds with agents responsible for isolated parts generated conflicts in interdependent elements

A sequential passage, with one person in charge of each coupled concern, produced greater progress.

Sharing helps.

Fragmenting without understanding the dependencies gets in the way.

How to test without turning the test into an endless project

Do not start with an entire campaign.

Choose a recurring, relevant, and easy-to-inspect asset. 

It can be a landing page, a weekly report, a sales sequence, or a campaign brief.

Then, run a governed pilot:

  1. Define the objective. Write the expected outcome without prescribing every detail of the execution.
  2. Fix the reference. Use a real example, an approved version, or a verifiable set of requirements.
  3. Write the rubric. Limit it to five or seven criteria that truly define quality.
  4. Separate constructor and critic. The producer should not be the sole source of evaluation.
  5. Set the limits. Establish cost, time, number of rounds, and situations that require human intervention.
  6. Choose the stopping criterion. Determine the minimum grade and which flaws prevent publication.
  7. Record the rounds. Keep version, review, change and result.
  8. Approve before publishing. The final decision needs an owner.
  9. Measure in the market. Choose one operational quality metric and one business metric.
  10. Update the pattern. What the market teaches must return to the next heading.

This sequence connects to a simple management logic:

  • Detect the main risk of the asset.
  • Direct with reference and criteria.
  • Execute with construction, criticism, and correction.
  • Grow using market feedback to improve the system.

Without this last step, the loop gets stuck inside the AI itself.

The criterion is the new infrastructure

Gauntlet Loop: AI without criteria just repeats the mistake faster

For a long time, companies treated quality as something an experienced person perceived when reviewing the work.

This used to work when the volume was lower.

With AI agents producing at scale, criteria need to move out of someone's head and into the system.

Not to eliminate human judgment.

To use it where it is most valuable.

The discussion about AI in companies is excessively focused on the model, the prompt, and the tool. But an operation doesn't become smarter just because it produces twenty versions instead of two.

She becomes smarter when she knows:

  • What you are trying to achieve.
  • How to recognize a good result.
  • Which limits cannot be exceeded.
  • When to stop the iteration.
  • What evidence should come from the market.
  • How to turn this evidence into a better pattern.

This also applies to B2B Growth Leadership

Leading is not overseeing every execution. 

It's keeping the system's telemetry clear enough so that speed doesn't destroy control.

Today, who defines what is good in your AI's work: a verifiable criterion or the agent itself that produced it?

Choose a single recurring asset and run the pilot. Do not increase autonomy before increasing the ability to evaluate.

If you'd like to receive practical insights on AI, acquisition, CRM, the sales pipeline, and revenue, subscribe to the Let's Test Brief.

If the problem is bigger than a single asset and lies in the connection between marketing, CRM, sales, and decision-making, a Growth Session It can help you organize the system before speeding up execution.

Share

Leave a Reply

Your email address will not be published. Required fields are marked *


Let's Test Brief 

A biweekly curation by Israel Degasperi on acquisition, CRM, the sales pipeline, and AI—all aimed at turning growth into revenue. With each issue, you’ll receive 1 central insight, 3 curated links, a practical application and the content that most highlighted on my channels.