Google's August 20 announcement introduced multi-campaign A/B testing for AI Max, complete with a demo showing ten clicks to launch an experiment. The trade press ran with the ease-of-use angle. What they missed: a faster setup does not mean a faster answer. If your hypothesis is weak, you will get a statistically significant result that tells you nothing about pipeline economics.
AI Max becomes the default for Search campaigns starting in September. Campaigns using automatically created assets and campaign-level broad match auto-upgrade first; Dynamic Search Ads follow in February 2027. The window for controlled testing is closing. The question is whether your test design will produce a decision you can defend in a pipeline review, or a dashboard number that looks good until someone asks about qualified lead rate.
The Experiment Architecture Google Built
AI Max experiments differ from traditional Google Ads experiments in one structural way that matters for B2B: they do not create a separate campaign copy. Google's documentation explains that the experiment splits your existing campaign 50/50 between control and treatment arms, keeping traffic and budget within a single campaign. This design accelerates learning by avoiding the cold-start problem of a duplicate campaign, but it also means you cannot inspect the treatment arm as an independent entity in your campaign view.
Joey Bidner's LinkedIn post captured the frustration many operators felt when they discovered this:
How are we supposed to properly validate these results if we cannot independently inspect what the treatment arm actually did?
Joey Bidner
Google's Ginny Marvin responded that the Experiments Results page does break out control and treatment arm results, but the point stands. You are trusting Google's reporting layer rather than verifying against a campaign you can audit directly.
The new capabilities announced in August address a different friction point. AI Max experiments now support brand and location controls, so you can test without disabling the guardrails that keep your campaigns from bidding on competitor terms or serving in geographies you do not cover. For B2B accounts with strict brand governance, this removes the excuse that testing AI Max required compromising compliance.
The Performance Claims and the Reality Gap
Google reports that AI Max campaigns see an average of 7% more conversions or conversion value at a similar CPA/ROAS when using the full feature suite compared to search term matching alone. Alphabet's Q2 2026 earnings disclosed that campaigns using AI Max and Performance Max together deliver an average of 15% more conversions at comparable return on ad spend.
Those are Google's numbers. Independent testing tells a different story. Digital Applied's analysis found that only 16% of advertisers in an independent poll reported good performance with AI Max, and Monks Agency testing showed 99% of AI Max impressions generated zero conversions across approximately 30,000 search terms. Leverage Marketing's tests found AI Max spent 20-30% of account budget to produce 10% of conversions, often at a higher CPA than traditional Search campaigns.
The gap between Google's case studies and typical advertiser experience is not a contradiction. It is a selection effect. The 7-15% lift figures come from accounts with clean conversion data, sufficient volume, and the patience to let the algorithm learn. The disappointing results come from accounts that turned on AI Max without fixing the inputs first.
Designing a Test Your CFO Will Sign
The ten-click setup is the easy part. The hard part is defining what success looks like before you launch. A B2B marketing executive running this test needs to answer three questions that Google's interface will not prompt you to consider.
First, what is your minimum detectable effect? If your campaign generates 50 conversions per month, you cannot reliably detect a 7% lift. The math requires either more volume or a longer test window. Running a two-week experiment on a low-volume campaign will produce a result, but the confidence interval will be wide enough to drive a truck through.

Second, what is your primary metric? Google's experiment reporting will show you conversions, conversion value, CPA, and ROAS. For B2B, none of these are the metric that matters. The metric that matters is qualified pipeline generated, and that data lives in your CRM, not in Google Ads. You need to tag experiment arm at the lead level and measure downstream qualification rate, not just form fills.
Third, what is your decision rule? Before you launch, write down the specific outcome that would cause you to roll out AI Max to the full campaign, the outcome that would cause you to reject it, and the outcome that would cause you to extend the test. If you cannot articulate these thresholds in advance, you will rationalize whatever result you get.
The Pilot Checklist
Adswerve recommends starting with no more than 10% of your search budget and allowing 8-12 weeks to see performance results. That timeline conflicts with the September auto-upgrade deadline for some campaign types, which means you may need to run a compressed test on campaigns that are not subject to auto-upgrade while letting the forced migrations happen on lower-priority campaigns.
Before launching, verify that your conversion tracking is complete. Browser-based tracking misses 30-60% of conversions, and AI Max optimizes on the signals you feed it. If your conversion data is incomplete, the algorithm will optimize toward the wrong thing with total confidence.
Set brand exclusions at the campaign level. Enable the brand and location controls that the August update made available. Document your negative keyword list and URL exclusions. AI Max's final URL expansion feature will send traffic to pages you did not specify unless you explicitly exclude them.
The Risk Nobody Wants to Model
The real risk of AI Max is not that it fails to deliver lift. The risk is that it delivers lift on the wrong metric. A 15% increase in form fills that produces a 20% decrease in qualified lead rate is a net loss for pipeline economics, but it will show up as a win in Google's experiment reporting.
PPC Geeks' analysis puts it directly:
The main risk is false positive uplift, where conversion volume rises but lead quality, margin or non-brand economics get worse.
PPC Geeks
Define your commercial pass marks before launching: CPA, ROAS, qualified lead rate, and margin thresholds. If you cannot measure lead quality at the experiment arm level, you are not running a test. You are running a hope.
The ten clicks are not the work. The work is the hypothesis, the measurement plan, and the decision rule. Get those right, and the experiment will tell you something useful. Get them wrong, and you will have a dashboard that looks great until the next pipeline review.