---
title: "Before an AI Agent Touches Your Google Ads, Run This 6-Week Holdout"
description: "Dedicated strategists run your account and a proprietary engine trained on $500B in profitable ad spend optimises the execution underneath them"
image: "https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6abdeef953a5c456c1a2f696_5c5ff0e3-0d2b-455b-a76d-57d4d7a9e4df.png"
---

October 1, 2026

•

11

min read

# Before an AI Agent Touches Your Google Ads, Run This 6-Week Holdout

![Young man with curly hair wearing a black shirt outdoors against green foliage background.](https://cdn.prod.website-files.com/6821efca072e48f6f495a47e/68562d390107b3921a6e3d68_1743932904108.jpg)

**Alexander Perleman**, Head Of Product @ groas
Ex-Goldman Sachs and Stanford Computer Science

**Email: alex@groas.com**

[**LinkedIn: https://www.linkedin.com/in/alexander-433793253/**](https://www.linkedin.com/in/alexander-433793253/)

![Cover image for: Before an AI Agent Touches Your Google Ads, Run This 6-Week Holdout](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6abdeef953a5c456c1a2f696_5c5ff0e3-0d2b-455b-a76d-57d4d7a9e4df.png)

A five-star review will not tell you whether an AI agent can beat your Google Ads manager. Neither will a one-star warning. On [Ryze AI’s Trustpilot page](https://www.trustpilot.com/review/get-ryze.ai), praise for automated budget shifts sits beside complaints about erratic bids and changes that never reached the account. Neither experience is your account, your conversion tracking, or your auction.

Give an [AI agent for Google Ads](https://emailgum.com/ryze-ai-review/) write access without a control group, and you are guessing. If revenue rises, the vendor takes credit. If it falls, auction competition takes the blame. I would run a concurrent holdout instead: same period, comparable traffic, one metric chosen before anyone sees a result. **Reviews tell you what happened to someone else. A controlled split tells you whether the agent beats your setup.**

#### Question: does the agent beat the current operator?

Write the hypothesis in the test brief before you grant access: **On comparable auction traffic, will autonomous management produce a lower cost per acquisition (CPA) or higher return on ad spend (ROAS) than the current operating model?** Choose *one* as the primary metric. Do not switch to the other because the first disappoints you.

This is not a contest to see who generates fifty responsive search ad variations fastest. Nor is it a test of whether a dashboard can produce more recommendations than a person can read. The agent claims it can manage the account more effectively. Measure the economic outcome of that management.

My expectation: continuous execution can beat periodic account reviews when it catches useful bid, budget, and query changes sooner. That is a mechanism, not a result. The split has to show whether the advantage survives contact with your account.

#### Setup: split traffic before anyone changes bids

Do not compare an agency-run May with an agent-run June. As PPC practitioner Tomas Kubilius explains in his guide to [clean Google Ads experiments](https://tomaskubilius.com/blog/clean-google-ads-experiments-setup-results/), a month-over-month comparison mixes the treatment with seasonality, supply-chain changes, and competitor bids. Run both operating models at the same time.

For standard Search campaigns, set up a **50/50 split in Google Ads Experiments**. Assign the current operator to the control and the agent to the treatment. Both sides then face the same test period and a split of eligible traffic. That does not make every individual auction identical; it removes the much larger problem of comparing different months.

![Diagram of a 50/50 ad traffic split into control and treatment arms with separate budgets and brand exclusions](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6abdeefa53a5c456c1a2f6cd_ea045c01-d681-4c50-bc9e-a2799344db67.png)

Where a suitable user-level split is not available for Shopping or Performance Max, use matched geographic markets. Pair territories using historical volume and conversion rate before assigning them to control or treatment. The draft pairing to investigate, for example, is Texas, Ohio, and North Carolina against Florida, Pennsylvania, and Georgia. Check the pairing against *your* history; state names alone do not make markets equivalent.

Do not give the agent high-intent core product campaigns while the incumbent gets expensive non-brand terms and call that a split. You have selected the winner before the test begins. **If you cannot make the arms comparable, do not treat the result as evidence of lift.**

##### Control: lock write access, not the incumbent’s hands

Document who can change what before day one. The incumbent should manage the control under its normal operating rules. The agent should manage only the treatment. Neither side gets to alter the other arm’s bids, negatives, assets, or landing pages. Keep conversion definitions and any sitewide changes consistent across both arms, and record exceptions.

This matters particularly when an agent connects through an API or Model Context Protocol connector. Operators discussing [business AI tools on Reddit](https://www.reddit.com/r/aiToolForBusiness/comments/1wrqgs0/ryze_ai_mcp_running_your_google_and_meta_ad/) call for explicit write boundaries, spend caps, and action logs. Put those boundaries in place before the connector goes live. A rollback for a serious problem is allowed; hiding it from the scorecard is not.

Give each arm its own hard daily budget cap. [Google Ads Help documentation](https://support.google.com/google-ads/answer/14147337?hl=en) says campaign experiments do not support shared budgets across arms. More broadly, a shared pool makes it harder to tell whether one manager performed better or merely had different access to spend. **Separate budgets and restricted write access are controls, not optional settings.**

##### Metric: pick the outcome and the win condition now

For lead generation, I would pre-register CPA on verified, non-spam conversions. For ecommerce, I would pre-register ROAS on net cart revenue after returns. Choose the measure that represents the business outcome, specify how it is calculated, and give both arms the same conversion source. Clicks, impression share, and ad relevance scores belong in the diagnostic notes, not in the victory column.

Set a practical hurdle before the first dollar is spent: for example, **15% lower CPA or 20% higher ROAS**, depending on which metric you chose. A 2% gap would not persuade me to replace an operating model. The hurdle is a business decision, though, not a statistical significance test. Record how you will assess uncertainty and the minimum conversion volume needed to make a call. Otherwise, someone will discover a new definition of “success” on day 43.

#### Sample check: can six weeks answer the question?

A calendar deadline cannot manufacture conversions. An account with eighteen conversions after four days cannot tell you much about a small CPA difference, however confident its dashboard looks.

[Konvtrack’s incrementality testing benchmarks](https://www.konvtrack.app/en/blog/incrementality-testing-google-ads-experiments) put the scale in perspective: its estimates call for approximately 1,500 conversions per arm to detect a 10% difference, or roughly 375 per arm for a 20% difference. Those figures are planning benchmarks, not a promise that your six-week test will reach either target. If the account gets fewer than 50 conversions a month across all campaigns, expect a multi-arm test to be inconclusive.

Check expected volume *before* the split. If the planned evaluation window will not produce enough evidence for the difference you care about, record that limitation. You can still learn whether the connector executes changes, respects budgets, and creates work for your team. You cannot turn a thin sample into proof of CPA or ROAS lift by speaking firmly about it.

##### Weeks 1–2: allow calibration, but keep the safety log

Treat the first two weeks as burn-in, not as part of the final performance score. [Google Ads API experiment guidance](https://developers.google.com/google-ads/api/docs/experiments/reporting) says to disregard the first one to two weeks when evaluating automated bidding models or new features. [Grow Wild Agency’s discussion of the learning phase](https://growwildagency.com/blog/google-ads-learning-phase-what-it-is-and-how-to-get-through-it/) describes roughly 50 conversions and 7 to 14 days per campaign for Smart Bidding calibration, with low-volume campaigns more prone to volatility.

That does not mean you ignore a runaway bid for fourteen days. Enforce the spend caps, inspect changes, and intervene if necessary. Log every intervention. **Weeks 1–2 are excluded from the outcome comparison, not exempt from supervision.** The planned scorecard covers weeks 3–6.

#### Weekly log: verify execution, not notifications

Every seven days, compare Google Ads Change History with the vendor’s account of what it did. If you are testing an [autonomous paid ads management tool](https://emailgum.com/ryze-ai-review/), look for changes that reached the account: bids adjusted, queries negated, assets paused or updated. A suggestion sitting in a notification feed is not an executed action.

![Paper ledger, inspection loupe, red checkmarks, and printed change logs on a desk](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6abdeefa53a5c456c1a2f6d0_3511626a-10c0-4733-8c27-792b8afb8af8.png)

This is the distinction I care about with [groas](https://groas.com/paid-search): its specialized models execute bids, budgets, and query filtering continuously, and log actions with plain-language reasoning. In a test, that should be visible in the platform record, not merely asserted in a sales deck. Keep three columns each week:

- **Actions executed:** What changed in Google Ads, cross-checked against Change History?
- **Performance observed:** What happened to conversion volume and the pre-registered metric for that cohort, allowing for conversion lag?
- **Human interventions:** What needed a rollback, repair, or manual decision, and how much operator time did it take?

Do not infer causation from a single weekly movement. The log is there to verify the mechanism and expose hidden labour. If an operator spends four hours a week repairing the agent’s decisions, put those hours beside the CPA or ROAS result. **Automation that quietly creates a second management job is not a clean win.**

#### Contamination check: three ways to break the holdout

Run this check every week. A tidy-looking 50/50 split is not enough if the treatment can take easier conversions or reach into the control.

1. **Brand query poaching.** If the agent or a Performance Max treatment bids freely on your brand while the comparison is supposed to measure generic acquisition, it can collect high-intent conversions and make its CPA look better. Search strategist Can Elmas’s breakdown of [bidding on your brand name](https://canelmas.com/blog/bidding-on-your-brand-name/) makes the case for isolating brand search and using negative brand lists across generic testing campaigns. Set those boundaries for both arms before launch. Do not let one side claim existing branded demand as new performance.
2. **Budget and audience leakage.** Shared budgets compromise the comparison. So does a geo design that allows the arms to reach overlapping target markets. Check campaign settings, location targeting, and spend pacing rather than assuming the labels “control” and “treatment” enforce themselves.
3. **Conversion lag distortion.** A recent week can look expensive because its conversions have not reached the CRM or posted back yet. Practitioners discussing [conversion lag on r/googleads](https://www.reddit.com/r/googleads/comments/1v8r3j8/conversion_lag_problem/) note that some high-ticket conversions take up to 30 days to mature. Plan a **14-day hold after week 6** before the first final read, then check it against your actual lag. If material conversions normally arrive later, wait for them before declaring a winner.

If one of these leaks occurs, document when it began and what it affected. Do not quietly clean the chart and call the experiment controlled.

#### Decision: win, loss, tie, or no answer yet

At the end of week 6, stop the evaluation window. After the planned maturation period, pull the same conversion cohorts and the same primary metric for both arms. Check volume, uncertainty, budget use, and the intervention log before making one of four calls:

- **Clear win:** The treatment clears the pre-registered CPA or ROAS hurdle on sufficient evidence, without requiring routine manual repair or sacrificing conversion volume to make the ratio look good. It has earned consideration for broader deployment.
- **Clear loss:** The treatment performs materially worse, or its operation requires unacceptable intervention. If CPA spikes or conversion volume collapses, use the stop rule and revert rather than paying for a cleaner-looking final chart.
- **Operational tie:** The primary outcome is close enough that the test does not establish performance lift, but the agent performs the work with less human effort. That can still favour autonomous execution over agency hours or manual account maintenance. Call it an operating-cost decision, not a statistically proven CPA victory.
- **Inconclusive:** Too few conversions, unresolved lag, or a contaminated split prevents a reliable comparison. “It’s still learning” is not a substitute for naming which of those problems remains and what another test would cost.

![Woodblock illustration of a balanced scale weighing CPA and ROAS markers beneath a plumb line](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6abdeefa53a5c456c1a2f6d3_240672da-61bf-4587-9ae5-484e00ffdaee.png)

If you are evaluating Ryze or [alternatives to Ryze](https://groas.com/post/best-ryze-alternatives-google-ads-2026-fully-autonomous-management), inspect what you are actually buying. Does the product execute within your guardrails, or does it hand you a queue of suggestions? A self-serve connector that leaves budget troubleshooting and negative-keyword cleanup with your team has not replaced the operator. It has assigned the operator homework. I would rather test autonomous execution with a named strategist accountable for direction than mistake a busy notification feed for management.

#### One-page brief: send this before granting access

Send any vendor or [AI Google Ads agency](https://groas.com/for-agencies) these terms before the product deck turns into a debate about features:

1. **Split:** Run a concurrent 50/50 Google Ads Experiment where suitable, or a pre-matched geographic split. No May-versus-June comparison.
2. **Boundaries:** Give each arm its own budget, write access, and brand-query rules. Log exceptions and interventions.
3. **Schedule:** Treat weeks 1–2 as burn-in and weeks 3–6 as the evaluation window. Enforce safety limits throughout.
4. **Success metric:** Choose verified CPA or net-revenue ROAS, set the hurdle and evidence requirement in advance, and record human management time separately.
5. **Final read:** Allow at least the planned 14-day conversion-lag buffer, extending it if your actual conversion cycle requires it.

A vendor willing to be measured should be able to discuss those boundaries. If an account executive insists on unrestricted brand access, offers only month-over-month summaries, or calls a control group unnecessary friction, decline the proposal. You would be funding the experiment while letting someone else write the answer.

If the agent clears the holdout, expand with hard budget guardrails and human ownership. Our runbook for [putting Google Ads on autopilot safely](https://groas.com/post/ppc-on-autopilot-how-to-run-google-ads-t) covers that next step. If it loses, keep the current setup. If the evidence is thin, do not pretend a star rating fills the gap. **Run the split, read the change log, and let the result change what you do next.**

#### Frequently asked questions

##### How can I tell if an AI Google Ads agent actually beats my current setup?

Run a concurrent holdout: put the current operator in a control arm and the agent in a treatment arm during the same period with comparable traffic. Reviews only show what happened to someone else's account, and comparing different months mixes seasonality and competitor changes into the result.

##### How do I split traffic fairly when testing an AI agent on Google Ads?

For standard Search campaigns, set up a 50/50 split in Google Ads Experiments with the current operator as control and the agent as treatment. For Shopping or Performance Max, where a user-level split may not be available, pair geographic markets using historical volume and conversion rate instead of comparing different months.

##### What boundaries should I set before letting an AI agent access my Google Ads account?

Give each arm its own hard daily budget cap and restrict write access so the agent manages only the treatment and neither side alters the other arm's bids, negatives, assets, or landing pages. Keep conversion definitions consistent across both arms and record every exception, including any rollbacks.

##### Which metric should I use to decide whether the AI agent won the test?

Choose one primary metric before spending starts: CPA on verified non-spam conversions for lead generation, or ROAS on net cart revenue after returns for ecommerce. Set a practical hurdle such as 15% lower CPA or 20% higher ROAS, and do not switch metrics if the first result disappoints.

##### How many conversions do I need for a Google Ads holdout test to be meaningful?

Planning benchmarks cited in the article estimate roughly 1,500 conversions per arm to detect a 10% difference and about 375 per arm for a 20% difference. If your account gets fewer than 50 conversions per month across all campaigns, expect a multi-arm test over six weeks to be inconclusive.

##### Why do the first two weeks of the test not count toward the result?

Google Ads API experiment guidance says to disregard the first one to two weeks when evaluating automated bidding models, so weeks 1–2 are treated as burn-in and excluded from the performance scorecard. The evaluation window covers weeks 3–6, followed by a 14-day hold after week 6 so lagging conversions can mature.

##### How do I verify the AI agent is actually making changes in my account?

Every seven days, compare the Google Ads Change History against the vendor's account of its actions, looking for executed changes like bid adjustments, negated queries, and paused or updated assets. A suggestion in a notification feed is not an executed action, so also log any rollbacks, repairs, or manual decisions and the operator time they took.

##### What are the main ways a Google Ads holdout test gets contaminated?

Three common leaks: brand query poaching, where one arm bids on your brand and claims high-intent conversions; budget or audience leakage from shared budgets or overlapping geo targeting; and conversion lag distortion, where recent weeks look expensive because conversions have not yet posted back. Isolate brand search, enforce separate budgets, and allow a 14-day post-test buffer before the final read.

## Related Posts

[![Cover image for: YouTube Ads Won’t Get Fewer: Six Calls to Check by 2027](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6abdf393d75bfe8c3da36a21_454d1aea-9ccc-43d4-977a-d7bf62eab14b.png) ##### YouTube Ads Won’t Get Fewer: Six Calls to Check by 2027 October 1, 2026 • 11 min read Written by David](https://groas.com/post/youtube-ads-won-t-get-fewer-6-prediction)

[![Cover image for: Your Google Ads Learning Phase May Take Months, Not a Week](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6abdeff30593563f84603596_4e800677-6b2b-4f54-98c1-2f2ccf7df073.png) ##### Your Google Ads Learning Phase May Take Months, Not a Week October 1, 2026 • 10 min read Written by David](https://groas.com/post/the-learning-phase-is-a-math-problem-how)

[![Cover image for: CPA, ROAS, and the Google Ads Bidding Terms People Keep Misreading](https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6abdefe635584b5e8b85f3f4_ad12c524-5634-4df5-ba58-621d946f5751.png) ##### CPA, ROAS, and the Google Ads Bidding Terms People Keep Misreading October 1, 2026 • 11 min read Written by David](https://groas.com/post/cpa-roas-and-the-bid-strategy-terms-adve)

![](https://cdn.prod.website-files.com/6821efca072e48f6f495a47e/6823c03839b2b618430d5ae6_Light-p-500.png)

[![White stylized owl eyes with green background icon.](https://cdn.prod.website-files.com/6821efca072e48f6f495a47e/6a68a036a5dd5627409d3579_Groas%20Icon.png)](https://groas.com/post/before-any-ai-agent-touches-your-google#)

[contact](mailto:scale@groas.ai)

Explore

[Home](https://groas.com/)[Philosophy](https://groas.com/our-philosophy)[Blog](https://groas.com/blog)

What We Do

[Paid Search](https://groas.com/paid-search)[Earned Search](https://groas.com/earned-search)

Who We Serve

[Businesses](https://groas.com/for-businesses)[Agencies](https://groas.com/for-agencies)

MCP

[ChatGPT](https://chatgpt.com/plugins/plugin_asdk_app_6a9c2d7108f08191beb5b3d736be8655?q=groas)[Claude](https://claude.ai/directory/mcp-groas-ai)

Legal

[Privacy Policy](https://groas.com/legal/privacy-policy)[Terms of Service](https://groas.com/legal/terms)

Get Started

[Apply](https://groas.typeform.com/to/xC1bQNUT)

© 2026 groas 🇺🇸

[![](https://cdn.prod.website-files.com/62434fa732124a0fb112aab4/62434fa732124a389912aad8_linkedin%20small.svg)](https://www.linkedin.com/company/groas/)

```json
{
  "@context": "https://schema.org",
  "@graph": [
    {
      "@type": "BlogPosting",
      "@id": "https://www.groas.com/post/before-any-ai-agent-touches-your-google#article",
      "headline": "Before an AI Agent Touches Your Google Ads, Run This 6-Week Holdout",
      "description": "",
      "url": "https://www.groas.com/post/before-any-ai-agent-touches-your-google",
      "image": {
        "@type": "ImageObject",
        "url": "https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6abdeef953a5c456c1a2f696_5c5ff0e3-0d2b-455b-a76d-57d4d7a9e4df.png"
      },
      "thumbnailUrl": "https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/6abdeef953a5c456c1a2f696_5c5ff0e3-0d2b-455b-a76d-57d4d7a9e4df.png",
      "datePublished": "2026-10-01T05:26:19.512Z",
      "dateModified": "2026-10-01T05:26:19.512Z",
      "inLanguage": "en-US",
      "isAccessibleForFree": true,
      "articleSection": "AI For Google Ads",
      "keywords": "Google Ads, AI For Google Ads",
      "mainEntityOfPage": {
        "@type": "WebPage",
        "@id": "https://www.groas.com/post/before-any-ai-agent-touches-your-google"
      },
      "author": { "@id": "https://www.groas.com/author/david#person" },
      "publisher": { "@id": "https://www.groas.com#organization" },
      "about": [
        { "@type": "Thing", "name": "Google Ads" },
        { "@type": "Thing", "name": "Pay-per-click advertising" },
        { "@type": "Thing", "name": "Performance marketing" }
      ],
      "mentions": [
        { "@type": "Organization", "name": "groas", "url": "https://www.groas.com" },
        { "@type": "Thing", "name": "Google Ads" }
      ],
      "speakable": {
        "@type": "SpeakableSpecification",
        "cssSelector": ["h1"]
      }
    },
    {
      "@type": "Person",
      "@id": "https://www.groas.com/author/david#person",
      "name": "David",
      "description": "",
      "jobTitle": "Founder &amp; CEO @ groas",
      "email": "",
      "url": "https://www.groas.com/author/david",
      "image": {
        "@type": "ImageObject",
        "url": "https://cdn.prod.website-files.com/6823bbd57170ea42b357cf81/686a2a76611e759fd8d8a3fd_David%20LinkedIn%20Profile.jpeg"
      },
      "sameAs": [
        ""
      ],
      "knowsAbout": [
        "Google Ads",
        "Performance Max",
        "AI Max",
        "Pay-per-click advertising",
        "Conversion tracking",
        "Bid management",
        "Search advertising"
      ],
      "worksFor": { "@id": "https://www.groas.com#organization" }
    },
    {
      "@type": "Organization",
      "@id": "https://www.groas.com#organization",
      "name": "groas",
      "alternateName": "groas.com",
      "url": "https://www.groas.com",
      "logo": {
        "@type": "ImageObject",
        "url": "https://cdn.prod.website-files.com/6821efca072e48f6f495a47e/6834a970e03a7016560e3515_Logo%20Design%20256x256.png"
      },
      "description": "Dedicated strategists run your account and a proprietary engine trained on $500B in profitable ad spend optimises the execution underneath them",
      "sameAs": [
        "https://www.linkedin.com/company/groas"
      ],
      "knowsAbout": [
        "Google Ads",
        "Performance Max",
        "AI Max for Search",
        "Pay-per-click advertising",
        "Autonomous campaign management",
        "AI advertising agents"
      ],
      "areaServed": "Worldwide"
    },
    {
      "@type": "BreadcrumbList",
      "itemListElement": [
        {
          "@type": "ListItem",
          "position": 1,
          "name": "Home",
          "item": "https://www.groas.com"
        },
        {
          "@type": "ListItem",
          "position": 2,
          "name": "Blog",
          "item": "https://www.groas.com/blog"
        },
        {
          "@type": "ListItem",
          "position": 3,
          "name": "AI For Google Ads",
          "item": "https://www.groas.com/category/AI For Google Ads"
        },
        {
          "@type": "ListItem",
          "position": 4,
          "name": "Before an AI Agent Touches Your Google Ads, Run This 6-Week Holdout",
          "item": "https://www.groas.com/post/before-any-ai-agent-touches-your-google"
        }
      ]
    }
  ]
}
```

```json
{"@context":"https://schema.org","@type":"FAQPage","mainEntity":[{"@type":"Question","name":"How can I tell if an AI Google Ads agent actually beats my current setup?","acceptedAnswer":{"@type":"Answer","text":"Run a concurrent holdout: put the current operator in a control arm and the agent in a treatment arm during the same period with comparable traffic. Reviews only show what happened to someone else's account, and comparing different months mixes seasonality and competitor changes into the result."}},{"@type":"Question","name":"How do I split traffic fairly when testing an AI agent on Google Ads?","acceptedAnswer":{"@type":"Answer","text":"For standard Search campaigns, set up a 50/50 split in Google Ads Experiments with the current operator as control and the agent as treatment. For Shopping or Performance Max, where a user-level split may not be available, pair geographic markets using historical volume and conversion rate instead of comparing different months."}},{"@type":"Question","name":"What boundaries should I set before letting an AI agent access my Google Ads account?","acceptedAnswer":{"@type":"Answer","text":"Give each arm its own hard daily budget cap and restrict write access so the agent manages only the treatment and neither side alters the other arm's bids, negatives, assets, or landing pages. Keep conversion definitions consistent across both arms and record every exception, including any rollbacks."}},{"@type":"Question","name":"Which metric should I use to decide whether the AI agent won the test?","acceptedAnswer":{"@type":"Answer","text":"Choose one primary metric before spending starts: CPA on verified non-spam conversions for lead generation, or ROAS on net cart revenue after returns for ecommerce. Set a practical hurdle such as 15% lower CPA or 20% higher ROAS, and do not switch metrics if the first result disappoints."}},{"@type":"Question","name":"How many conversions do I need for a Google Ads holdout test to be meaningful?","acceptedAnswer":{"@type":"Answer","text":"Planning benchmarks cited in the article estimate roughly 1,500 conversions per arm to detect a 10% difference and about 375 per arm for a 20% difference. If your account gets fewer than 50 conversions per month across all campaigns, expect a multi-arm test over six weeks to be inconclusive."}},{"@type":"Question","name":"Why do the first two weeks of the test not count toward the result?","acceptedAnswer":{"@type":"Answer","text":"Google Ads API experiment guidance says to disregard the first one to two weeks when evaluating automated bidding models, so weeks 1–2 are treated as burn-in and excluded from the performance scorecard. The evaluation window covers weeks 3–6, followed by a 14-day hold after week 6 so lagging conversions can mature."}},{"@type":"Question","name":"How do I verify the AI agent is actually making changes in my account?","acceptedAnswer":{"@type":"Answer","text":"Every seven days, compare the Google Ads Change History against the vendor's account of its actions, looking for executed changes like bid adjustments, negated queries, and paused or updated assets. A suggestion in a notification feed is not an executed action, so also log any rollbacks, repairs, or manual decisions and the operator time they took."}},{"@type":"Question","name":"What are the main ways a Google Ads holdout test gets contaminated?","acceptedAnswer":{"@type":"Answer","text":"Three common leaks: brand query poaching, where one arm bids on your brand and claims high-intent conversions; budget or audience leakage from shared budgets or overlapping geo targeting; and conversion lag distortion, where recent weeks look expensive because conversions have not yet posted back. Isolate brand search, enforce separate budgets, and allow a 14-day post-test buffer before the final read."}}]}
```
