Skip to main content

Google Ads

Google Ads Ad Strength 'Excellent' converted at 1.3% vs 3.8% for 'Good' across $75k

7 min read

Every responsive search ad in Google Ads carries a label: Poor, Average, Good, or Excellent. It sits in the ads table, it shows up in recommendations, and it's the first thing a new account manager is told to fix. "Get all your ads to Good or Excellent" is close to universal advice.

We pulled ten days of performance for every enabled ad in a mid-sized search account (roughly $75,000 of spend), and grouped it by that label. Here is what came back.

Ad Strength

Spend

Impressions

Clicks

Conversions

CTR

CVR

Cost per lead

Excellent

$9,548

12,385

890

11.2

7.19%

1.3%

$850

Good

$6,787

8,331

481

18.5

5.77%

3.8%

$367

Average

$42,111

67,237

3,767

114.0

5.60%

3.0%

$369

Poor

$16,759

12,666

1,196

35.0

9.44%

2.9%

$479

Ads rated Excellent converted at a third the rate of ads rated Good, and cost 2.3× more per lead. Ads rated Poor had the highest click-through rate in the account, which is the one thing Ad Strength is most often assumed to predict.

If the label meant what people think it means, that table would be upside down.

What Ad Strength actually measures

Google's documentation is reasonably candid about this, but the candour rarely survives into agency playbooks. Ad Strength scores four things:

Factor

What it's checking

What it rewards

Headline count

How many of the 15 headline slots are filled

More headlines

Description count

How many of the 4 description slots are filled

More descriptions

Keyword inclusion

Whether the ad group's keywords appear in the copy

Keyword insertion or literal repetition

Asset uniqueness

How different the headlines and descriptions are from each other

Variety

Note what is not on that list: click-through rate, conversion rate, cost per conversion, quality score, landing page match, anything downstream of the impression. Ad Strength is calculated before the ad has served a single impression. It cannot be a performance metric because there is no performance in it.

It is a score for how much raw material you have handed the combination engine. That's all. Google's own help centre says it "is not a factor in Ad Rank" and has no direct bearing on auction outcomes.

Why more material can convert worse

Here is the mechanism, and it's not mysterious.

A responsive search ad with 15 unique headlines and 4 descriptions can be assembled into thousands of combinations. Google picks the combination per auction. The more diverse your assets (the more Ad Strength rewards you), the more combinations exist that no human ever wrote or reviewed.

Ad design

Headlines

Combinations possible

Message consistency

Tight, 3 pinned headlines

3

1

Every impression says the same thing

Moderate, 6–8 headlines, some pinned

8

dozens

Mostly on-message

"Excellent", 15 diverse headlines, nothing pinned

15

thousands

Whichever three the algorithm chose

The tight ad says one thing, and the landing page says the same thing, and the searcher who clicked knows what they're going to get. The Excellent ad says one of several thousand things, and some of those combinations promise something the landing page doesn't deliver, and the searcher who clicked on "Free Quote Today" and lands on a rebate explainer bounces.

This is the ad-to-landing-page consistency argument, and it is a well-established driver of conversion rate. Ad Strength rewards the thing that works against it.

There's a second mechanism, visible in the CTR column. Poor-rated ads (fewer, more specific headlines), had the best click-through rate. That's consistent with specificity winning the click: "6.6kW Solar From $X" beats a rotating selection of generic benefit statements for someone who searched for a 6.6kW system. Diversity is Ad Strength's proxy for relevance. It is not the same thing.

The honest caveat

This is correlational, and the correlation is confounded in at least three ways.

Confound

Why it matters

Ad Strength is not randomly assigned

High-strength ads tend to be in ad groups someone recently worked on, which may differ in keyword mix, bids, and landing page

Volume is uneven

Excellent has 890 clicks and 11 conversions; Average has 3,767 and 114. The Excellent bucket has wide confidence intervals

Campaign mix

If Excellent ads cluster in a campaign that converts badly for unrelated reasons, the label takes the blame

So the finding is not "Excellent ads convert worse." The finding is "in this account, across $75,000, the label did not predict conversion rate in the direction everyone assumes, and the ads it rated worst had the best CTR." That is enough to stop treating the label as a KPI. It is not enough to conclude the label is inversely predictive everywhere.

The rigorous version is an experiment. Google Ads has one built in: an ad variation, run inside a single ad group, comparing an Excellent-rated ad against an Average-rated one with otherwise identical targeting. That removes every confound in the table above at once. We have the harness in place and will report the result.

What would change our mind

Evidence

Would suggest

Ad variation test shows Excellent variant converting better

The label has predictive value once confounds are removed

Excellent ads' CVR rises as their volume grows

The 1.3% is small-sample noise

Pinned-headline Excellent ads outperform unpinned ones

It's the unpinned diversity, not the label, that hurts

Same pattern repeats across three or more accounts

It's structural, not account-specific

Until one of those lands, the burden of proof belongs on the score, not on the sceptic.

Where this leaves the "get everything to Excellent" advice

It's advice about inputs masquerading as advice about outcomes. Filling all 15 headline slots is easy to do, easy to verify, and easy to report to a client as progress. None of those properties make it correlated with cost per lead.

What the advice optimises

What the client wants

Number of assets

Number of leads

Asset diversity

Message clarity

A green label in the interface

A lower number on the invoice

The two columns overlap sometimes. They are not the same, and in this account they pointed in opposite directions.

What we recommend instead

Measure the label against your own outcomes before you optimise toward it. Ads report → add the Ad Strength column → add Conversion rate and Cost per conversion → sort. Two minutes. If your Excellent ads aren't outperforming, you have your answer for this account, and no general claim about Ad Strength overrides it.

Pin what matters. A pinned headline in position 1 guarantees every impression leads with your core message. Ad Strength penalises pinning. Ignore that.

Write fewer, sharper headlines. Eight good headlines with two pinned will usually beat fifteen diverse ones. The label will drop to Average or Good. That is fine.

Run the test. One ad group, two ads, one variable. If Excellent wins, use it. If it doesn't, you've learned something about your own account that no help-centre article can tell you.

Never put Ad Strength on a client report as a KPI. It is a checklist item at most. Reporting it as a performance metric teaches the client to value something that, on the evidence, doesn't predict what they're paying for.

A related finding: identical ads, different CTR

While this account was running a landing page test, it produced a data point that bears directly on what Ad Strength can and can't predict.

The test used ad variations: the same responsive search ads, byte-for-byte identical headlines and descriptions, with only the final URL changed. On the first full day both arms served:

Test

Control CTR

Variant CTR

Gap

Battery (6 campaigns)

3.16%

6.35%

+101%

Panel (1 campaign)

3.60%

5.53%

+54%

Same copy. Same Ad Strength label, necessarily. A 54–101% difference in click-through rate, because the variant ads were new, and Google's rotation gives new creative an exploration boost while it estimates their expected CTR. The gap decayed over the following days as real data replaced the estimate.

The point for Ad Strength: CTR is driven substantially by serving decisions, not just by the copy. Two ads with identical text and identical Ad Strength scores produced wildly different CTRs because of when they were created. A label computed from the text alone can't see any of that, which is one more reason it doesn't predict the outcome.

What a test-ready ad group looks like

The account's highest-volume battery campaign had three enabled ad groups. Each carried exactly one responsive search ad.

Ad group

Enabled RSAs

Ad Strength

Home Battery

1

Average

Solar Battery

1

Average

Solar & Battery Package

1

Poor

One ad per ad group means there is nothing for the rotation to choose between and nothing to compare. It's also below Google's recommended minimum of two RSAs per ad group. But the fix is not "get them to Excellent": it's a second ad in each group, written to test one specific hypothesis against the first, so that in a month there's a result rather than a label.

The wider point

Google Ads surfaces a growing number of scores: Ad Strength, Optimization Score, Quality Score, the various "recommendation" percentages. Some of them are genuinely useful. Some are inputs dressed as outcomes. The way to tell the difference is the same every time: pull the score alongside your conversion data and see if it predicts anything.

For Ad Strength, in this account, it didn't. Check yours.

Related services

Filed under

google-adsad-strengthresponsive-search-adsconversion-ratersaad-copy

Keep reading

More from Insights