
Open any responsive search ad in your account and you’ll see it: a little strip of colored bars, a label that reads Poor, Average, Good, or Excellent, and a nagging feeling that you should probably do something about it if it’s not green. Google puts this indicator right where you’re writing your ads, which means it’s the very last thing you look at before you publish. That placement alone convinces a lot of advertisers that Ad Strength is something worth optimizing for its own sake — a target on par with Quality Score or conversion rate.
It isn’t, and the gap between what advertisers assume Ad Strength means and what it actually measures is wide enough that it’s worth walking through carefully. This post covers what Ad Strength is actually built from, what Google itself says it does and doesn’t affect, what happens when you look at real performance data instead of the label, and what to do with that information without swinging to the opposite extreme of ignoring it entirely.
What Ad Strength actually measures
Ad Strength is a rating Google Ads generates for a responsive search ad while you’re building or editing it. It’s calculated from three things: relevance, quantity, and diversity of the assets you’ve written — the headlines and descriptions that make up the ad.
Relevance looks at whether your headlines reflect the keywords in the ad group, including variations of them, not just a single repeated phrase. Quantity is close to a checklist: Google wants you using close to the maximum number of assets, which for an RSA means fifteen headlines and four descriptions, and the rating drops if you’re only filling in the minimum. Diversity is about whether those fifteen headlines actually say fifteen different things — a different feature, a different benefit, a different angle — rather than fifteen near-duplicates of the same value proposition with the words shuffled around.
Put those three together and you get a five-point scale: Incomplete, Poor, Average, Good, and Excellent. Incomplete usually just means you haven’t filled the ad out enough for Google to judge it. The other four are where the actual grading happens, and where most of the anxiety lives, because Google surfaces this label prominently in the ad editor and in account-level recommendations, often alongside a suggestion to add more headlines or rewrite ones it considers too similar to each other.
None of this is secret or in dispute. It’s the mechanical part of Ad Strength, and Google is reasonably transparent about how the score is built. The part that causes confusion is what people assume that score then does.
You’ll run into the same rating in two places, and they don’t always agree. Inside the ad editor, Ad Strength is calculated live as you type, updating headline by headline, which is what makes it feel like immediate feedback on quality rather than a periodic audit. At the account level, Google Ads also surfaces Ad Strength as a recommendation category — a running tally of how many of your ads sit at each label, usually paired with a suggested action like “improve ad strength for 6 ads” sitting in your Recommendations tab next to bid and budget suggestions. Treating that recommendation with the same skepticism you’d apply to any other auto-generated suggestion in that tab is worth doing, because it’s competing for your attention against changes that do have a measurable effect on spend and clicks.
What Ad Strength is not measuring
Here’s the sentence worth remembering: Ad Strength does not feed into Ad Rank, it is not a component of Quality Score, and it plays no role in whether your ad wins an auction. It is a content-completeness and variety check, run at write-time, meant to nudge you toward practices Google’s own research suggests correlate with better outcomes on average. That’s a fundamentally different thing from a ranking signal.
This distinction matters because Quality Score already exists, already gets calculated per-auction from expected click-through rate, ad relevance, and landing page experience, and already has a well-understood relationship to what you pay and how often you show up. If you haven’t looked at how that one actually works, it’s worth reading our breakdown of what Quality Score actually measures and how to improve it, because the two ratings sit right next to each other in the interface and get conflated constantly. Quality Score has teeth — it genuinely affects your cost per click and your position in the auction. Ad Strength has none. It’s advisory, evaluated once when you save the ad, not recalculated per auction, and disconnected from Ad Rank entirely.
Google has said as much directly, and third-party analysts who’ve dug into the mechanics agree: Ad Strength is meant to help you build a better ad during the creation process, not to predict how that ad will actually perform once it’s live. The label you see is a snapshot of compliance with Google’s own best-practice checklist, not a forecast.
The data: what happens when you actually test the theory
Best-practice checklists are usually harmless even when they’re a little reductive. The problem shows up when you check whether following the checklist actually correlates with the results advertisers care about — and when someone has actually run that check at scale, the answer is uncomfortable.
Optmyzr, a PPC management and reporting platform, published a study analyzing more than a million ads across responsive search ads, expanded text ads, and Demand Gen creative, pulled from over 22,000 Google Ads accounts that had been active for at least 90 days and spending a minimum of $1,500 a month. That’s a large enough sample, and a strict enough filter for real, ongoing accounts rather than brand-new or dormant ones, that the results are worth taking seriously rather than dismissing as an outlier.
The headline finding: Ad Strength does not reliably correlate with performance. Ads labeled Poor were not, on average, the worst performers in the dataset. They were, in fact, the best. Poor-rated ads delivered the highest average ROAS of any label in the study — 327.65% — outperforming Average, Good, and Excellent-rated ads on that metric. If your mental model says an Excellent label should translate to your best-performing ad, the data says otherwise, at least in aggregate across this sample.
The study also dug into one of the specific mechanics Google’s scoring rewards — asset quantity, which in practice pushes advertisers toward writing longer headlines to fill out the character limit and hit that fifteen-headline target with substantial-sounding copy. Here the data cuts the other way just as clearly. Headlines under 20 characters produced a cost per acquisition of $9.35, against $18.27 for longer headlines — nearly half. Short headlines also won on click-through rate (11.77% versus 10.52%), conversion rate (10.39% versus 8.61%), and conversions per impression (1.22% versus 0.91%). Every metric in that comparison points the same direction, and it’s the opposite direction from what “write more, write fuller headlines to maximize your score” would suggest.

Pinning headlines — locking a specific headline into a specific position instead of letting Google’s system choose the combination — is another area where Ad Strength penalizes you for a decision the data doesn’t clearly punish. Google’s own guidance treats pinning as something that limits the algorithm’s flexibility and can suppress your rating. The study found some pinning was, in fact, the strongest approach by cost per acquisition, return on ad spend, and cost per click, with no pinning at all a close second on CPA. The tradeoff was conversion rate, which did suffer under pinning — a genuinely useful, specific finding, but a more nuanced one than “never pin or your score will drop,” which is what the interface implies every time you lock a headline in place.
None of this means Ad Strength is actively harmful advice, and none of it means you should race out and write the shortest, most repetitive ads you can. What it means is that the score is measuring compliance with a set of authoring conventions, and those conventions are, at best, weakly and inconsistently related to the outcomes that actually matter for your account.
Why the mismatch happens
It’s worth understanding the mechanism, not just the data, because “the score is wrong” isn’t quite the right takeaway either. Ad Strength is doing exactly what it was designed to do — it’s just measuring something narrower than advertisers assume.
The scoring logic rewards things a machine can check mechanically: are there fifteen headlines, do they use different words, do those words include keyword variants. It has no way to evaluate whether a headline is actually persuasive, whether it matches the specific intent behind a specific search, or whether your fourth headline is a genuinely distinct value proposition versus filler copy written purely to hit the asset count. An account manager writing to maximize Ad Strength is, in practice, often writing to satisfy a content-diversity checker, not to persuade a human being who just typed a search query.
There’s also a simpler, more mundane reason the mismatch persists: Ad Strength is evaluated once, at write time, against the ad in isolation. It has no visibility into your actual auction results, your landing page, your offer, or how your ad compares to the three competitors showing up in the same auction. A headline can be rated highly diverse and still lose to a competitor’s ad that says something more compelling in fewer words, because diversity within your own set of fifteen headlines says nothing about how any of them stack up against what else is on the results page. Performance is inherently relative and contextual; Ad Strength is neither.
This is the same failure mode that shows up whenever a platform turns a soft, judgment-based practice into a hard, gradeable score. The advice underneath — use varied language, cover different angles, don’t leave asset slots empty — is genuinely reasonable as a starting checklist for someone new to writing RSAs. The problem is that a checklist score gets gamed the moment it becomes a target in itself, and “write fifteen sufficiently different-sounding headlines” is a much easier bar to clear by padding than by writing fifteen headlines that would each independently earn a click.
Ad Strength is not the only Google score worth this kind of skepticism
If this pattern sounds familiar, it’s because Google Ads has more than one built-in score that nudges you toward its own definition of a well-run account, and Ad Strength isn’t even the most aggressive of them. Optimization Score is the more visible example — a percentage Google puts directly on your dashboard, tied to a running list of recommendations, most of which involve giving Google’s automated systems more control or more budget.
We’ve written before about what to actually apply from your Optimization Score, and why chasing 100% can quietly cost you, and the underlying lesson applies just as well here: a score that Google generates, displays prominently, and frames as something to maximize is not automatically a proxy for your account’s actual performance. Sometimes it’s genuinely useful signal. Sometimes it’s a nudge toward the outcome that benefits Google’s own systems and revenue as much as it benefits you. The only way to tell the difference is to check the score against your own results rather than assuming alignment.
Ad Strength sits in a slightly different category than Optimization Score, because it isn’t asking you to hand over budget or bidding control — it’s asking you to write ads a particular way. But the underlying discipline is identical: treat the label as one input worth glancing at, not a scoreboard, and always let your actual conversion and cost data have the final word.
So should you ignore it completely?
No, and this is where a lot of the commentary on Ad Strength overcorrects. The score is a poor predictor of performance, but the underlying checklist it enforces isn’t useless — it’s just being asked to do a job it can’t do, which is stand in for persuasion quality.
Where Ad Strength earns its keep is at the floor, not the ceiling. An ad rated Incomplete usually means you’ve genuinely left value on the table — too few headlines, missing descriptions, an ad that hasn’t given Google’s system enough raw material to test combinations against. Catching that early, before an ad group launches with three headlines instead of ten, is a legitimate and useful thing for an automated check to flag. The same goes for genuinely duplicate headlines that add nothing — if three of your fifteen headlines are functionally identical, that’s wasted inventory in the rotation regardless of what any study says about ROAS by label.
Where it stops being useful is at the top end, where the difference between a Good and an Excellent rating usually comes down to hitting an arbitrary asset count or rephrasing a headline to look more different from its neighbors on paper, neither of which has much to do with whether the ad converts. Treat the label as a compass pointing toward “did I do the basic setup work,” not a destination you’re trying to reach for its own sake.
It helps to think of the scale as having a useful half and a mostly cosmetic half. Incomplete and Poor are worth acting on quickly, because they usually flag a real gap — too few assets, or assets so similar to each other that Smart Bidding has almost nothing to test against. Average, Good, and Excellent are a much softer gradient, and the honest answer is that the practical difference between them, once you’ve cleared the Poor threshold with genuinely varied copy, is small enough that the study data above couldn’t find a consistent performance pattern across them at all. Spend your editing time closing the gap from Poor to Average with real content. Don’t spend it closing the gap from Good to Excellent by padding.
A practical framework for writing RSAs that actually convert
Given the data, here’s a more useful way to approach responsive search ads than chasing the green label:
- Write for length discipline, not headline count for its own sake. Short, specific headlines outperformed longer ones on every metric in the Optmyzr data. Don’t pad a headline to fill the character limit if the shorter version says the same thing more clearly.
- Fill all fifteen headline slots, but make each one earn its place. The quantity guidance isn’t wrong on its own — more genuine variations give Smart Bidding more combinations to learn from. The mistake is filling slots with filler just to clear Incomplete or Average. If you can’t think of a fifteenth genuinely different angle, a strong twelve beats a padded fifteen.
- Build real diversity around angles, not synonyms. Price, speed, guarantee, social proof, specific feature, use case, objection-handling, urgency — that’s a list of genuinely different angles. Rewriting the same value proposition four ways with different adjectives is what Ad Strength rewards and what the performance data doesn’t.
- Use pinning deliberately, not to appease the score. If there’s a headline that must appear — a legal disclaimer, a specific brand name, an offer that has to be in position one for compliance reasons — pin it and accept the Ad Strength hit. The data suggests some pinning can still perform well on cost metrics; what it costs you is some conversion rate, which is a real tradeoff to weigh, not an automatic red flag.
- Check the search term and asset performance reports before rewriting for the label. Google Ads shows you which individual headlines and descriptions are pulling their weight under each ad’s asset detail view. That’s real performance data on your own account, and it should outrank a generic score every time the two disagree. If you haven’t built the habit of checking search terms regularly, our guide to reading a Google Ads search term report is a good companion practice — the same instinct of trusting your own data over the platform’s summary applies to both.
- Don’t let Ad Strength override a specific business reason for shorter or narrower copy. A regulated industry, a highly specific niche product, or a brand voice that doesn’t lend itself to fifteen distinct angles are all legitimate reasons to stay below Excellent. A lower label with copy that’s accurate and on-brand beats a higher label achieved by stretching claims to hit an asset count.
Where this gets more complicated: AI Max, automatic assets, and 2026’s push toward AI-generated copy
This debate is getting more urgent, not less, because of where Google is pushing responsive search ads next. At Google Marketing Live 2026, Google announced Real-Time Policy Reviews for RSAs, cutting the ad review process from hours down to seconds, and continued expanding Asset Studio’s ability to generate headlines, descriptions, and other creative assets automatically using generative AI. Combine that with campaigns increasingly running under AI Max — which layers automatic asset generation and broader query matching on top of standard Search campaigns — and you get a system that can fill out all fifteen headline slots with plausible-sounding, superficially diverse copy in seconds, with no human ever weighing in on whether any individual line is actually persuasive or accurate.
We’ve covered the mechanics of that shift in more depth in our post on what AI Max for Search actually changes and what you still control, but the relevant point for Ad Strength specifically is this: an AI system optimizing to satisfy Google’s own asset-diversity checker is going to be very good at generating an Excellent-rated ad, because that’s a well-defined, gradeable target. It has no particular reason to be good at writing copy that reflects your actual differentiators, matches your compliance requirements, or avoids the kind of industry-specific overreach that shows up when generative text doesn’t understand the nuance of a regulated category. As more of the asset-writing gets automated, the gap between “scored highly” and “actually correct for this business” doesn’t close. If anything, it’s easier than ever to end up with a wall of Excellent-rated ads that nobody has actually read.
That argues for the opposite instinct from the one the interface encourages. The more automated the asset generation gets, the more valuable it becomes to have a human check the actual headlines being served — not just the score sitting next to them — on some regular cadence.
How to use Ad Strength without being used by it
A short mental checklist to run against the label each time it shows up:
- Is this ad missing headlines or descriptions outright? If yes, that’s a real gap — fix it regardless of what the study data says about labels.
- Are any of the headlines functionally duplicate — same claim, different words? Trim those regardless of the score.
- Is the score suggesting I lengthen a headline that’s already clear and specific? Ignore it; shorter tends to win.
- Is the score penalizing a pin I have a real business reason for? Keep the pin, note the tradeoff, move on.
- Have I actually looked at this ad’s real CTR, conversion rate, and cost data recently, independent of its label? If not, that’s the number that should be driving the next edit, not the color of the bar.
Applied consistently, this turns Ad Strength back into what it was supposed to be in the first place — a basic completeness check you glance at once while building an ad — rather than a recurring source of second-guessing every time you open the editor.
It’s also worth setting a review cadence rather than reacting to the label every time it changes. Ad Strength can shift after Google’s own automated text customization edits a headline, or after a competitor’s changes alter what “diverse” looks like relative to the auction — none of which reflects anything you did. Checking asset-level performance monthly, alongside whatever cadence you already use for search terms and budgets, is enough to catch a genuine problem without turning every dashboard refresh into a reason to rewrite ads that are already working.
A few questions worth settling
Does a low Ad Strength rating hurt my Quality Score? Not directly. Quality Score is calculated from expected CTR, ad relevance, and landing page experience for the specific query that triggered the ad — Ad Strength isn’t one of its inputs. There can be an indirect relationship if a genuinely thin, low-effort ad also happens to perform poorly enough to drag down expected CTR over time, but that’s a real-performance effect, not the label itself doing anything.
Will Google restrict my ad’s reach if Ad Strength is Poor? No. A Poor-rated ad still enters every auction it’s eligible for and competes on the same Ad Rank basis as an Excellent-rated one. The label affects what Google recommends you change, not what the ad is allowed to do.
Should new advertisers ignore Ad Strength from day one? Not entirely — for someone who’s never written an RSA before, the underlying checklist (fill the slots, don’t duplicate headlines, cover a few different angles) is a reasonable starting structure. The advice in this piece isn’t “the checklist is worthless,” it’s “don’t keep optimizing past the point where the checklist stops correlating with results,” which for most well-built ads happens well before you’d need to hit Excellent.
Does this apply to Performance Max asset groups too? Performance Max uses a similar asset-strength concept for its asset groups, built on the same relevance-quantity-diversity logic, and the same caution applies — it’s a completeness signal, not a performance guarantee. If you’re auditing a Performance Max campaign more broadly, our piece on how to audit a Performance Max campaign when you can’t see inside it covers the wider set of things worth checking beyond any single asset score.
Why this is easier to catch with eyes on the account regularly
The theme running through all of this isn’t really about Ad Strength specifically — it’s about the gap between a platform’s own scorecards and what’s actually happening in your account, and how quickly that gap can widen when nobody’s checking. An Excellent-rated ad that’s quietly underperforming its Poor-rated neighbor isn’t a dramatic failure. It’s the kind of small, easy-to-miss mismatch that sits there for months if the only thing anyone glances at is the color of the label, because on the surface everything looks fine.
That’s the specific gap Growera’s continuous daily account monitoring is built to close — not by replacing the judgment calls in this post, since deciding whether a shorter headline or a pinned asset is right for your business is exactly the kind of call that needs a person weighing your specific tradeoffs, but by making sure the actual performance numbers behind every ad get looked at regularly enough that a scorecard mismatch gets caught in weeks, not discovered by accident during an annual audit. If you’re curious how that continuous check-in works alongside the rest of an account, the Growera homepage walks through the full picture, and the pricing page lays out what’s included at each tier.
The bottom line
Ad Strength measures whether your responsive search ad clears a content checklist — enough headlines, enough descriptions, enough variety between them. It does not measure, predict, or guarantee ad performance, and Google itself is explicit that it plays no role in Ad Rank, Quality Score, or auction outcomes. The most rigorous look at real account data so far, a study of over a million ads across 22,000-plus accounts, found no reliable correlation between the label and results, and found Poor-rated ads posting the best average ROAS of any category, with shorter headlines beating longer ones on cost, click-through rate, and conversion rate across the board.
None of that makes the label worthless. It’s a reasonable floor check — catch genuinely thin ads, catch actual duplicate headlines, make sure you haven’t left slots empty. It stops being useful the moment you start rewriting good copy to satisfy an asset-diversity checker instead of a customer, and that’s exactly the point at which most advertisers currently are chasing it. Write for the person reading the ad, check your own conversion and cost data before you check the label, use pinning when you have a real reason to, and treat “Excellent” as a nice-to-have footnote rather than the goal. That’s a more useful relationship with Ad Strength than either obsessing over it or dismissing it outright — and it’s the same discipline worth applying to every score Google Ads hands you that isn’t actually your bottom line.
Ready to see ERA in action?
Connect your Google Ads account and see what it finds — free for 14 days.
Start your free trialNo card charged until day 14 · cancel anytime