Ad Testing at Scale, and Why Ours Sometimes Says Nothing

The short version
- Ad copy testing works in a small account and stops working in a large one. The arithmetic breaks before the discipline does.
- Our system reads every ad group, finds the weakest ad in each, says why, and writes replacement variants for that specific ad group.
- In ad groups without enough volume to judge, it returns nothing at all. That restraint is the part worth writing about.
- On one account, monthly leads fell 55% while the leads that met the client’s hiring requirement held steady, taking the qualifying rate from 28% to 65%.
- That account does not sell anything. The advertising exists to staff the business, which makes it an operations function rather than a sales one.
- Hiring did not fall with the volume. 352 leads produced 31 hires in the spring. 162 leads produced 29 in the summer.
- Cost per conversion over the last 28 days is down a third, from $125.56 to $84.57.
Testing ad copy is the oldest discipline in paid search and the first one to get dropped. Not because anyone decides to drop it. Because of how the arithmetic works.
In an account with twenty ad groups you can look at every ad, review performance, form a view, and write something better. At two hundred it becomes a monthly spreadsheet exercise that slips whenever anything else in the account needs attention. At two thousand there is no version of the week in which a person gets through them all. So third-party software gets introduced, or testing narrows to the priority ad groups: the big spenders, the ones a client asked about, the ones that broke recently.
Everything else keeps running whatever was written when it was built. Often years ago, often by somebody who has since moved on.
Why the obvious fix does not work
The obvious fix is to sort by spend and work down the list until the week runs out.
It helps, and it misses the same thing a spend-sorted search terms report misses. The ad groups at the top are the ones already getting attention. The waste is distributed: a few hundred ad groups each performing slightly below what they could, none of them individually large enough to reach the top of anybody’s list, and collectively the largest number in the account.
Sorting by spend finds the ad groups you were going to look at anyway.
What we built
Each run reads every ad group in the account and, within each one, compares the ads against each other. Not against a benchmark and not against the account average, because an ad group’s ads compete for the same impressions in the same auctions, and the only fair comparison is the one against the ads standing next to it.
Where there is a clear weakest ad, the system returns four things:
- which ad, by ad ID, ad group, and campaign so changes can be made in bulk
- why it is the weakest, in a sentence a person can read and disagree with
- replacement variants, written for that ad group rather than for the account generally
- what the variants are trying to fix, so the next run has something to measure against
Those three identifiers together — ad ID, ad group, campaign — decide whether any of this gets used. A recommendation that a widely used variant is underperforming is easy to action at scale, and that is exactly the risk: changed account-wide, it also strips the variant out of the ad groups where it was working. The granularity is what lets you fix the ad groups that need it and leave the rest alone.
Automating that recurring audit across the whole account’s ad creative, identifying improvements at the level of the individual ad group, and delivering them in a bulk-import format is the difference between proactive performance marketing and account maintenance.
The part most of this depends on
Both this system and the search term and negative keyword work run against the same thing: a written description of the business, what it sells, who it sells to, and what it will not say. For ad copy that last part carries most of the weight. A regulated industry has claims it cannot make, an employer has requirements it must state, and a generated headline that ignores either is worse than no headline.
That description is not a setting somebody ticks. It is written with the client, it is specific, and it is the reason the output reads like the business rather than like an ad generator. Neither system is better than the description it is given.
The part that matters: sometimes there is no answer
Every ad group gets a minimum volume check before it gets a verdict. Below it, the system returns nothing for that ad group. No weakest ad, no variants, no recommendation.
This is the feature, and it is the one worth explaining.
In an ad group with eleven clicks spread across three ads, one ad has the lowest click-through rate. It always does. Something has to be last. But that ranking is noise wearing a rank, and acting on it means rewriting an ad for a reason that will reverse itself next week. Do that across a few hundred ad groups and the account is full of changes made for no reason, with a testing record that cannot distinguish the changes that mattered from the ones that were coin flips.
A system built to produce answers will always produce one, because producing one is what it was built for. Deciding that the honest output is “not enough data to say” is a choice somebody has to make deliberately, and it is the choice that makes every other answer worth acting on. If the system declines to guess where it cannot know, then when it does name a weakest ad, that means something.
It is the same principle behind labeling a change “not separable from noise” rather than reporting it as movement, which is described in how we measure. A number nobody should act on is worse than no number, because a number gets acted on.
What happened on a real account
Marketing is usually judged on what it sells. This account sells nothing.
The client advertises to hire. The campaigns produce no revenue of their own and are measured on one thing: whether the business can staff itself fast enough to take on the work in front of it. That makes the advertising an operations function rather than a sales one, and it changes the definition of a good lead. Not somebody who might buy. Somebody the business can actually employ.
It is worth saying plainly because it is the half of this work that gets least attention. Growth is not only a demand problem. A business that cannot hire cannot take the work it has already won, and advertising is as capable of solving that constraint as it is of creating demand.
The role carries a requirement that most people searching the general job title do not meet. Someone who does not meet it is not a near miss. They are someone the business cannot hire, however keen they are and however well the ad performed.
So this account has an unusually clean quality measure. Every lead either meets the requirement or it does not, and the ratio between them is a number that cannot be argued with.
What went wrong first
In the spring, Google broadened how it matched the keywords this account bids on. Searches from people looking for the far more common version of the role, which had historically stayed outside the account’s targeting, started arriving in volume.
The bidding strategy made it worse rather than better, and it did so by working correctly. Bidding was optimizing toward conversion value weighted for qualified applicants, which is the right configuration. But conversion value is learned from what converts, and a flood of unqualified applicants who happily fill in a form looks like conversion volume. The account was buying more of exactly the traffic it could not use.
In April, 159 leads produced 44 that met the requirement. A 28% qualifying rate.
What the two systems did
The search term and negative keyword automation started reading every term weekly and judging it against the business description rather than against a word list, which is what it takes to separate the searches the account wants from the ones naming the same job and meaning something it cannot hire. From August the ad performance system ran alongside it, finding the underperforming ad in each ad group and generating replacement copy written against the same description.
What happened
| Month | Leads | Met the requirement | Qualifying rate | |—|—:|—:|—:| | April | 159 | 44 | 28% | | May | 92 | 54 | 59% | | June | 101 | 55 | 54% | | July | 91 | 51 | 56% | | August | 71 | 46 | 65% |
Total leads fell 55%. The leads that could actually be hired did not fall at all, going from 44 to 46. The qualifying rate more than doubled.
The number that matters to the business is further down the funnel, and it moved the right way too. In April through June, 352 leads produced 31 hires. In July and August, 162 leads produced 29. Less than half the volume, effectively the same hiring. Hiring is lumpy and it lags, so the month-to-month figures swing more than the quarters do, and we would not read a single month of it either way.
On the advertising account over the last 28 days, compared with the 28 before: conversions rose from 18 to 41 and cost per conversion fell from $125.56 to $84.57, a reduction of roughly a third, on spend that rose from $2.26K to $3.47K. September is running at a 73% qualifying rate, though that month is still in progress and we are not counting it yet.
What we cannot tell you
We cannot separate the two systems after August. The negative keyword work was done in the spring, and the account’s change history shows why the early move belongs to it: 252 of the 253 negative keywords ever added went in during March and April. So the jump from 28% to 59% in May is that work, not this one. The ad copy system only started running in August, and the improvement from that point is the two of them together. Any split we offered for that period would be invented.
That is an unsatisfying sentence in a piece about one of the two, and it is the only accurate one. An account is not a laboratory. Both systems were deployed because the account needed both, not in an order designed to make attribution easy, and a vendor who hands you a clean percentage attributable to each of two simultaneous changes is telling you something they cannot know.
Two other limits, stated plainly. The lead and hire figures come from the client’s recruiting records and the cost figures come from the advertising account, so the two sets of numbers count different things and will not reconcile against each other. And the qualifying rate is unusually measurable here because the hiring requirement is binary. Most businesses do not have a quality measure that clean, which is a fact about this account rather than a claim about the method.
What it does not do
It does not change anything in the account. Every variant is a candidate, and a person approves it.
That is deliberate at this size. The risk with an automated system holding write access is not one bad change, it is a great many of them made before anybody looks, in an account too large to review after the fact. Keeping a person in the approval step costs a few minutes a week and removes that whole category of problem.
It also does not tell you whether the offer is right, whether the landing page matches the promise, or whether the ad group should exist at all. Those are judgment questions about the business, and the system has no view on them. What it does is make sure every ad group gets looked at in a week, which is a question of time rather than a question of judgment.
And it will sometimes be wrong about why an ad underperformed. That is why the reason is written out, and why a person reads it before anything changes.
Where the scale claim comes from
The same system runs on an account spending $5 million a month across thousands of ad groups. That account is our founder’s own venture, where the benefits of this change the bottom line of the entire business.
It is also the hardest test of the restraint described above. At that size a false positive is not one bad recommendation, it is a great many of them arriving every week — which is where a system that declines to guess earns its keep.
It runs on your account
This plugs into the Google Ads account you already have. There is nothing to install on your website, no platform to move onto, and no rebuild involved. The account is the only thing it needs.
Where a site we built adds something is after the click. When the site, the analytics and the CRM are wired together from the start, a change to an ad can be read against what it produced rather than stopping at the account boundary, so ad copy, conversion work and media budget are judged against the same data. That is an advantage on top, not a condition of entry.
Both systems run on client accounts today alongside the rest of the Google Ads management work. If your account has more ad groups than anyone can read in a week, tell us what you are working with.