Your $4,000 agency retainer cleared again. Your non-brand CPA did not move. Ask whether the agency is earning its fee and you may get a slide deck; freeze a baseline and watch the next 30 days, and you get something you can argue with.
That is the test. Not whether the report looks polished, but what changed, who changed it, and what happened to acquisition cost per dollar of fee. Pull a clean 90-day baseline, inspect the platform’s change history, track qualified pipeline, and compare the next month with what came before. A before-and-after audit is not a causal holdout experiment. It is a much better starting point than letting blended CPA and a few green arrows settle the invoice.
The question: what is the fee buying?
A performance retainer should buy deliberate execution, not just reporting. Agency retainers commonly run 10% to 20% of ad spend, or $1,500 to upwards of $10,000 a month. The pitch is familiar: specialists will examine search queries, add negatives, test creative, adjust budgets, and protect margins. Google’s automated bidding can handle much of the mechanical work. That does not make strategy worthless. It makes the work behind the fee worth inspecting.
Start with the attribution problem. Brand searches and warm audiences can make an account look healthy while non-brand acquisition struggles. A true holdout test addresses the causal question: how many conversions would have happened without the advertising? This protocol does not answer that question on its own. It does expose whether the agency is doing identifiable work and whether performance improves against a consistent baseline. Otherwise, vanity metrics can make a stalled account look healthy.
Write the question at the top of the spreadsheet: What did this month’s management fee buy that I can see in execution and qualified results?
Setup: freeze the account’s starting line on day 0
Export 90 days of performance, then separate brand
In Google Ads, export the trailing 90 days through yesterday. Use daily campaign-level data segmented by campaign type and conversion action. Keep an untouched copy. Record total ad spend, conversion rate, pipeline value if available, and non-brand cost per acquisition (CPA).
Pull campaigns containing your trade name into a separate sheet. If the business is called Apex Plumbing, searches for Apex Plumbing do not belong in the same acquisition calculation as searches for an unfamiliar plumber. Check generic Search and Performance Max reporting for brand traffic too; a campaign name alone does not tell you what demand it captured.
Divide non-brand spend by non-brand primary conversions. Label the result Baseline_CPA. If your CRM tracks qualified pipeline, calculate Baseline_Cost_Per_SQL the same way, using sales-qualified leads rather than all conversions. Keep the definitions fixed for the next 30 days. Change the denominator halfway through and you have tested your spreadsheet skills, not your agency.
Export change history and identify the actor
Open Google Ads Change History and export the trailing 30 days. Google Ads records account changes in Change History; use that platform record rather than a retrospective list assembled for a meeting. Sort by User and distinguish named team logins from system or API changes. Record what changed, when, and by whom.
Do not turn the count into a verdict yet. Ten useful negatives may matter more than 30 cosmetic edits, and an automated system can perform valuable work without a person clicking through the interface. An almost empty history is a question your agency needs to answer: where is the execution, and how can you verify it?

Match ad conversions to qualified pipeline
Now check the conversion actions used for bidding. Separate primary conversions from secondary diagnostics, then match Google Ads conversions against your CRM or lead-tracking sheet. Keep three buckets distinct:
- Raw inquiries: Form submissions and inbound calls lasting over 30 seconds.
- Sales-qualified leads (SQLs): Inquiries that meet your buyer criteria and schedule an evaluation call or consultation.
- Closed-won revenue: Booked contracts or completed ecommerce purchases.
A PDF download is not a sale. An unverified form fill is not automatically a qualified lead. If bidding optimizes for easy micro-actions, Google can get very good at finding people who take those actions. Calculate the share of non-brand inquiries that become SQLs and preserve it as your pipeline-quality baseline. A falling reported CPA is not a win if fewer leads qualify.
Observation window: collect four weekly snapshots
For the next 30 days, leave the campaigns alone. Do not make your own edits and do not change the agency’s instructions mid-test. You do not need to announce the audit; you do need to note any outside change that could affect the comparison, such as a new offer or landing page. Otherwise, you will not know what moved the result.
Every Monday, export the trailing seven days of Change History and the matching performance data. Use the same date convention each week. Keep the exports, not just a running tally. Your notebook should answer four questions.
1. What actions happened, and who made them?
Group changes by named users, system activity, and API activity. For each meaningful action, note its purpose: negative-keyword pruning, bid-target adjustment, budget reallocation, creative test, or conversion-tracking change. Look past the total. Four changes that address a costly search-term pattern can be more useful than 40 toggles with no clear reason.
When an account shows little activity, expect a familiar explanation: touching it would disturb machine learning. That may explain a particular decision to leave a setting alone. It does not explain a month with no visible review, no search-term decisions, and no account of what the team chose not to change. Other advertisers have raised the same change-history question and debated what an empty log means. Your job is not to impose an edit quota. It is to demand an account of the work.
2. Did search intent improve?
Review the Search Terms report alongside the change log. Identify irrelevant queries, negatives added, bidding targets revised, creative tested, and budgets moved between product lines. Broad and phrase match can reach useful variations; they can also spend against intent you would never choose deliberately.
Ask your agency which terms it reviewed, which it excluded, and why. A search-term review should leave a decision trail, even when the decision is to keep a term. If most recorded activity consists of accepting auto-apply recommendations, ask what independent judgment the retainer purchased.

3. What happened to non-brand CPA and pipeline?
At day 30, calculate non-brand CPA using the same conversion actions and brand exclusions as your baseline. Compare conversion volume, spend, and the share of inquiries that reached qualified pipeline. Check year-over-year search trends or industry indexes before treating every change as agency-driven; demand shifts can flatter or punish a month’s results.
If you need a causal estimate rather than a directional audit, investigate a formal Conversion Lift holdout experiment for Search or Performance Max where your account is eligible. Independent geo-based incrementality testing illustrates why platform-reported return and incremental return are not interchangeable. Do not label a simple before-and-after CPA movement incremental lift.
4. What did the fee cost per additional conversion?
Put media and management fees on the same page. Suppose you spend $15,000 on ads and pay a $4,000 retainer. At 50 non-brand conversions, platform CPA is $300; all-in acquisition cost is $380 ($19,000 divided by 50). The fee is not free just because the dashboard omits it.
Next, compare conversion volume with the frozen baseline, adjusted as fairly as you can for spend and demand. If that comparison suggests 40 conversions at the prior run rate and the test month delivers 50, you have 10 additional observed conversions. Dividing the $4,000 fee by 10 gives a $400 fee per additional conversion. It does not prove the agency caused all 10. It does tell you what the invoice must justify. When the difference is small or uncertain, say so rather than printing a heroic incrementality claim.
Read the result without worshipping the edit count
Set decision rules before you see the final month. Mine would be demanding, not universal. They are designed for an established account with enough activity to make a monthly review worth doing:
- Execution: The agency can show deliberate, relevant work through Change History and explain important decisions that do not appear there. For a manually managed account, fewer than 10 meaningful changes in a month triggers a direct explanation, not an automatic conviction. If execution is automated, ask for its action trail instead.
- Non-brand performance: Aim for at least a 15% reduction in non-brand CPA against the baseline, or 20% more primary conversions at or below Baseline_CPA. Check spend, demand, and conversion definitions before calling either outcome a pass.
- Pipeline quality: The ratio of CRM-verified SQLs to raw inquiries holds steady or improves. More cheap forms with fewer qualified buyers is a fail, whatever the deck says.
- Fee alignment: Compare the retainer with the gross margin associated with additional conversion volume. My preferred ceiling is 20%, provided you can make a credible estimate of that additional volume.
Spend growth deserves its own inspection. If budget rises from $10,000 to $20,000 while non-brand CPA stays flat or worsens, a percentage-of-spend fee can rise even though your acquisition efficiency does not. That is the incentive problem behind percentage-of-spend pricing. If spend climbs 30% and non-brand CPA drifts upward, ask what the added budget bought before approving another increase.

If the agency clears your pre-set rules, show it the scorecard and keep the relationship. If it misses most of them, do not wait for another polished PDF to explain away the same result. Ask for a corrective plan tied to the baseline, then decide whether the retainer still earns its place.
Turn your baseline into a written target
A failed audit often sends owners toward CPA-priced or revenue-share providers. Read those offers just as closely. Paying per lead can reward loose form validation and piles of unverified inquiries. Paying a share of gross revenue can become expensive as sales grow. A different invoice formula is not automatically better alignment.
Contract against the frozen baseline and a defined outcome, with clear conversion definitions, brand exclusions, guardrails, and exit rights. groas describes its 90-day growth sprints as written performance commitments, such as +30% more sales or +40% search visibility, backed by a commitment to work for free until the target is hit. Whatever a provider promises, make it specify the starting number and the measurement method before work begins.
Put these questions in front of any CPA-priced, pipeline-priced, or autonomous alternative:
- How do you define and audit the baseline? Require a written treatment of branded traffic and CRM-verified pipeline. Do not let gross conversions that include existing demand stand in for new acquisition.
- How is execution handled each day? Ask for an action trail, not a claim of constant attention. At groas, specialized models handle bidding, negatives, and budget reallocations within client guardrails while a named strategist owns direction and accountability. Judge that model, too, by the actions and outcomes you can inspect.
- What makes the provider more money? A percentage of spend rewards a bigger budget; a share of revenue can compress margins as volume rises. Ask for the fee, target, and cancellation terms in writing.
Who should skip the 30-day test?
If you spend under $2,000 a month on Google Ads, a single month may produce too few conversions for a useful CPA comparison. Do not manufacture activity by demanding daily edits. Skip the protocol for now if you launched a new site, changed an unproven offer, or entered a new market within the last 60 days. You need a usable baseline before you can test movement against it.
This is for an established account spending roughly $5,000 to $50,000 or more a month, with performance that has been flat for months and a retainer that keeps clearing. Freeze the 90 days. Collect the four weekly snapshots. Price the fee against qualified results, not report length. If the agency passes, you have a stronger reason to keep paying it. If it fails, the baseline is the number to put in front of whoever wants the account next. Make them put their target beside it in writing.

