Google Ads experiments help you compare a proposed campaign change with a control under a defined test setup. Before changing the campaign, write down the decision you need to make, the outcome that will decide it, the budget you can risk and the evidence you will accept.
That preparation matters because a campaign can produce more conversions while sending less useful enquiries to sales. A lower cost per conversion can accompany a fall in total qualified demand. An encouraging first week can also change when delayed conversions arrive. The experiment needs a business question as well as a platform configuration.
Our guide to improving PPC advertising ROI introduces testing as part of campaign improvement. This guide turns that idea into a practical test plan: select a supported workflow, define the comparison, monitor its health and make a decision that respects uncertainty.
The planning examples and outcome figures in this guide are fictional. They demonstrate decisions; they do not describe completed Canada Create experiments or results from a real advertising account.
Start with the decision the experiment must support
A useful question sounds like, “Should we add broad match versions of these keywords to this eligible Search campaign?” or, “Should this Performance Max asset group use the proposed creative set?” Each question points to a specific action that can be accepted, rejected or deferred.
“Which campaign is best?” is harder to answer because “best” could mean more sales, lower acquisition cost, greater reach or less work for the team. Narrow the decision before selecting a success metric. Write the campaign, audience, offer and outcome in ordinary language so that the person approving spend can understand what the result would change.
Identify the consequence of a wrong decision. Replacing a low-risk message may need a different evidence standard from moving a major revenue stream to a new campaign type. Record implementation cost, reversibility and how much deterioration the business could tolerate while learning. These considerations inform the plan; they do not turn weak evidence into strong evidence.
Also distinguish a problem that needs repair from a question that needs an experiment. A broken enquiry form, an incorrect phone number or duplicate conversion recording needs investigation and correction. Keeping a known failure in a control group to create a dramatic result wastes the opportunity to test a useful commercial hypothesis.
Finally, name the decision owner. Marketing may operate the campaign, sales may define qualification and finance may assess profitability. The brief should identify who resolves disagreements and who can approve adoption. Without that ownership, a completed test can remain an unresolved argument about metrics.
Match the question to the supported experiment route
Google Ads experiments are a family of workflows. Campaign eligibility, tested settings, allocation and reporting differ. The following capabilities were checked against official Google documentation on October 8, 2026. Confirm the current options for the actual campaign before committing spend.
| Campaign or question | Documented route and distinction |
|---|---|
| Search and Display | Custom experiments compare supported campaign changes, including bidding and targeting. Search offers cookie-based and search-based assignment; Display uses cookie-based assignment. See custom experiment setup. |
| Search ad messaging | Ad variations support changes to responsive search ad text and URLs across selected ads or campaigns. See ad variation setup. |
| Search features | Dedicated AI Max experiments compare AI Max off and on within one campaign. New broad match experiments compare original keywords with original keywords plus broad match versions. These are different treatments. New broad match requires supported Smart Bidding and excludes shared budgets, portfolio bidding and active trials, among other restrictions. |
| Performance Max | Google distinguishes uplift, upgrade and optimization experiments. Adding Performance Max alongside other campaigns asks a different question from replacing Standard Shopping with Performance Max. See Performance Max experiment types. |
| Demand Gen | Dedicated A/B experiments support creative, audience, feed and bidding comparisons. Google advises against testing budget as the variable. In its asset A/B route, control changes sync to treatment except budget adjustments. See Demand Gen experiments. |
| Video | The basic video asset A/B workflow supports Video Reach campaigns using Efficient Reach and Video View campaigns. Google directs other video campaign types to Custom. Its dedicated guide describes Custom experiments with up to ten arms. See video experiment setup. |
| Standard Shopping | The documented one-click Target ROAS experiment compares Target ROAS with manual CPC in eligible Standard Shopping campaigns. This differs from a Shopping-versus-Performance-Max test. See Shopping Target ROAS experiments. |
| App | Dedicated asset workflows include uplift and directional experiments. Both App uplift and directional experiments currently support Android only. Directional experiments are limited to video-only App campaigns for installs. See App asset setup, uplift eligibility and directional experiments. |
| Hotel Ads | Google’s custom setup instructions also list Hotel Ads, with specific bidding restrictions. Review the Hotel Ads notes in the setup guide. |
We build and run Google Ads campaigns measured on calls, forms and booked jobs, not clicks.
Within Performance Max, testing the addition of video differs from comparing two creative sets. Google documents asset addition tests and a separate asset-set A/B workflow labelled Beta. The latter compares sets within one asset group; common assets can continue serving alongside each assigned set.
Describe a feature bundle honestly. A Performance Max text customization and Final URL expansion experiment evaluates the documented combination. It does not isolate the effect of a single sentence on one landing page. Similarly, testing AI Max as a package cannot establish which component produced the observed difference.
Do not assume that a missing campaign can be made eligible by removing important settings. Investigate the reason and reassess the plan. Changing campaign structure just to unlock a test may create a different question from the one the business intended to answer.
Write a hypothesis that can be contradicted
A hypothesis connects the proposed change to an expected outcome through a plausible reason. Use this structure: “For this campaign and offer, changing X should improve Y because Z.” Then identify evidence that would make you reject the hypothesis.
For a fictional equipment supplier, a hypothesis might be: “Adding broad match versions of our existing keywords will find more qualified enquiries within our agreed acquisition-cost limit because buyers describe the same equipment need in ways our current keyword set may miss.” The expectation concerns qualified demand, not simply additional searches.
Separate the observation from the explanation. The supplier may have evidence that useful customer language appears in sales notes. That supports investigating reach. It does not prove broad match will find those buyers economically. The experiment tests that proposition.
Specify what remains constant: offer, service area, sales response process, conversion definitions and any settings outside the treatment. Some platform workflows intentionally change several elements together. Where that happens, name the package and limit the conclusion to that package.
Write the competing explanation too. Broader reach might bring more unsuitable enquiries; a new creative might attract curiosity without purchase intent. These possibilities help select guardrails and diagnostic metrics. They should not become excuses invented after an unfavourable result.
For a bidding comparison, avoid changing the landing page or ads at the same time. Google’s bidding test guidance recommends keeping the comparison focused. A clear treatment makes the eventual decision easier to defend, even when the answer is to keep the current approach.
Define the primary outcome before choosing report columns
The primary outcome is the measure that will decide whether the proposed change meets the test’s purpose. Choose one primary outcome, then use a small number of secondary measures to explain what happened and guardrails to protect the business.
A lead-generation test might prioritize qualified-enquiry volume subject to a cost limit. Another might prioritize cost per qualified enquiry while protecting volume. These are different decision rules. The first asks whether additional demand is affordable; the second asks whether efficiency improves without losing too much demand.
Define a qualified enquiry in observable terms. For the fictional equipment supplier, it could mean an identifiable organization requesting an eligible product in the service area, with a usable contact route and a confirmed commercial requirement. The definition must be applied consistently to both arms.
Record the exact conversion action and data source. If the experiment’s conversion column combines form submissions, phone clicks and qualified leads, its cost per conversion is not automatically cost per qualified lead. A similar label does not make two metrics equivalent.
When the business outcome lives in the CRM, verify before launch that the planned analysis can associate outcomes with the appropriate experimental groups and preserve the relevant timing. Do not assume every experiment workflow exposes the identifiers or grouping needed for an external analysis. If it cannot support the intended measurement, revise the design rather than attaching an unmeasurable business promise to it.
Keep exploratory observations separate. A change in device mix or search terms may explain a result and suggest the next test. It does not replace the primary outcome because that outcome was disappointing. Google’s significance assessment for one reported metric cannot be transferred to a different CRM metric that has not received its own suitable analysis.
For a deeper investigation of unsuitable enquiries and sales feedback, use our Google Ads lead quality guide. Resolve that diagnosis before treating a new campaign setting as the answer.
Freeze the definition before launch. If sales starts using a stricter qualification rule halfway through, record the break and reassess comparability. The same principle applies to conversion windows, value definitions and changes to what the campaign optimizes toward.
Decide whether this campaign can answer the question
Eligibility is the first check. Feasibility is the next. A campaign can be eligible for a platform experiment and still provide too little information to resolve the business decision within an affordable period.
Review historical volume for the actual primary outcome, its variability, typical spending and the time outcomes take to mature. A large number of clicks does not solve a shortage of qualified enquiries. Equally, a campaign that changes sharply with weather, availability or a small number of large orders needs a plan that acknowledges that variation.
Define the smallest improvement worth acting on before considering duration. This is a commercial threshold, not a statistical confidence level. An improvement that saves less than the cost of implementation and oversight may be measurable without being useful. A larger improvement may justify action, but the test still needs enough information to distinguish it from normal variation.
Google’s Campaign Guidance describes how volume, variability, allocation, duration and experiment type affect power. The documented scope is Performance Max experiments and broad match experiments in Search. Its estimate can help compare eligible setups. Google also notes that an estimate based on historical data can differ from the actual experiment.
Use that guidance as an input, not a promise. Ask whether the campaign is likely to produce enough relevant information within the spending limit and business season. An experiment planned around a year-end promotion may not answer how the same setting behaves in ordinary trading.
If feasibility is weak, consider a more informative campaign, fewer variants, a longer affordable window or a more focused question. Do not quietly replace a qualified-lead outcome with page views just to make the report populate. That would answer a different question.
Sometimes the appropriate decision is to defer. Improve the measurement foundation or accumulate stable operating history, then revisit the experiment. Record the reason: insufficient eligible volume, unresolved tracking, an imminent offer change or a budget that cannot sustain a useful comparison. Deferral is a planning outcome, not a failed campaign.
Specify assignment, traffic and budgets separately
Allocation determines which opportunities belong to each arm. Spending determines the resources available to pursue them. They are related, but the same percentage does not necessarily describe both.
For eligible Search custom experiments, cookie-based assignment uses cookies to keep users in an arm, while search-based assignment occurs for each search. A person making repeated searches may encounter both versions with search-based assignment. These options do not amount to a universal guarantee about identifying a person across every device or browser.
Google recommends a 50% split in its custom experiment instructions. Use the options supported by the selected workflow. A more cautious treatment allocation can limit exposure to a risky change, but it may leave less information in that arm. Decide the tradeoff before launch.
A traffic split does not guarantee equal realized spending, clicks or conversions. Google’s experiment FAQ distinguishes budget division from division of eligible traffic and describes scaling in unequal-split reports. Record whether each reported total is raw or normalized, and which comparison basis the selected workflow uses. Do not apply a second adjustment to an already normalized result.
For example, an arm assigned less traffic could produce fewer total leads while having a better reported rate. That does not justify comparing the two raw lead counts as if exposure had been equal. Conversely, an apparent efficiency improvement does not make the absolute volume irrelevant to staffing or revenue needs.
List every participating campaign and its budget in the brief. Calculate the expected total exposure to spend across the test, including any added treatment campaign. Record an authorized financial limit and the review process for staying within it. A limit written in a document is not an automated platform spending control.
Check competing activity before launch. Promotions, overlapping tests and changes to sales coverage can complicate interpretation. You do not need to freeze the entire business, but you do need to know which external changes would undermine the decision or restrict its scope.
Complete the experiment brief
Download the experiment briefs, deviation logs and decision records (ZIP).
A brief should be short enough to review together and specific enough to prevent disagreement later. Use the following fields as a reusable template. Replace every placeholder before scheduling; an empty approval, budget or measurement field is an unresolved dependency.
| Field | Required entry |
|---|---|
| Decision and owner | Whether to adopt the named change in the named campaign; decision owner and rollback owner. |
| Hypothesis | Expected outcome, proposed mechanism and evidence supporting the question. |
| Platform route | Campaign type, experiment subtype, eligibility evidence and campaign identifiers recorded privately. |
| Control and treatment | Exact settings, assets or feature package; what remains consistent. |
| Primary outcome | Metric, conversion action, source, direction and qualification rules. |
| Commercial threshold | Smallest worthwhile improvement and its economic or operating rationale. |
| Allocation and spend | Assignment method, split, participating budgets, authorized exposure and monitoring owner. |
| Guardrails | Limits for quality, volume, spend, tracking and customer-facing failures. |
| Timing | Start, learning allowance where applicable, analysis window, lag rule and final review date. |
| Evidence standard | Selected report confidence level and method for any separately analysed business outcome. |
| Stopping and adoption | Rules for harm, success, inconclusive results, extensions, auto-apply and manual implementation. |
Attach only evidence needed to operate the plan: a baseline summary, measurement definition, settings comparison and relevant eligibility notes. A large export that nobody can interpret is less useful than a concise record of the assumptions being tested.
Keep customer records and credentials out of broadly shared planning documents. Reference the authorized evidence location and responsible owner. The brief needs traceability without becoming a second uncontrolled copy of the CRM.
Worked planning example: decide whether broader Search reach is worth testing
This is a fictional planning exercise. No account has been inspected, approved or changed. Alder Equipment Supply sells commercial equipment within a defined service region. Its team wants more qualified enquiries and proposes adding broad match versions of existing keywords.
The first draft says, “Increase conversions.” During planning, the sales owner points out that some form submissions concern unsupported products. The team revises the primary outcome to qualified enquiries and defines the product and geography requirements before discussing a success threshold.
The proposed control retains the existing keyword set. The proposed treatment adds broad match versions through an eligible documented workflow. Bidding, ads, destination pages, offer and qualification rules remain outside the proposed treatment. The team must confirm actual eligibility and available reporting before selecting the final route.
For illustration, the business would consider additional qualified demand worthwhile only if acquisition cost remains within its documented economics and sales can respond within normal working hours. Those are required inputs, not numbers borrowed from another advertiser. Until finance and sales supply them, the example remains a draft brief.
The team chooses qualified-enquiry volume as the primary outcome, with cost per qualified enquiry and invalid-enquiry workload as guardrails. This differs from making acquisition cost the primary outcome. The objective is useful growth within limits, so a small cost increase could be acceptable if it stays within the predefined ceiling and creates worthwhile additional demand.
The measurement owner must establish whether qualified outcomes can be analysed for the experiment’s groups. If only raw form submissions are observable, the team cannot honestly claim to have planned a qualified-demand test. It must resolve that gap or explicitly narrow the question.
Before launch, the team also checks whether a scheduled product promotion will overlap the evaluation window. If it does, the plan can either target performance during that promotion or choose a more representative period. Calling the result “normal trading performance” would be unjustified if the evidence came entirely from a special offer.
The resulting decision is conditional: proceed only after eligibility, measurement, economics, timing and spending authority are resolved. That is the value of the brief. It exposes the missing pieces while changing them is still inexpensive.
Translate the brief into a setup review
Use the brief beside the current Google Ads setup screen. Confirm that the selected campaign and experiment subtype match the intended question. Review the original campaign’s status, supported settings and any warning that affects eligibility. Follow the dedicated guide for the chosen route rather than adapting steps from a different campaign type.
Compare the proposed arms before scheduling. Check settings, conversion goals, assets, destination URLs and approved ads. If an asset cannot serve, the experiment may evaluate a different creative package from the one described in the brief. Resolve that discrepancy or document the narrower treatment before starting.
Record how synchronization works. For custom Search and Display experiments, sync is on by default and cannot be changed after the trial is created. It copies base-campaign changes to the trial, not the reverse. Shared-ad edits affect both arms even with sync off. Synced changes are not currently reflected in Change history, so keep a separate dated record. These rules do not describe every experiment route.
Review automated actions too. Recommendations, scripts, scheduled promotions and other operators can alter the experiment without appearing in the brief. Assign one person to coordinate campaign changes and keep a dated deviation log. This is an operating agreement, not a reason to disable necessary business controls without review.
Confirm the end-state behaviour. Google’s monitoring guide says auto-apply is on by default for the experiment types that support it. Check the selected route and actual setting before launch. If adoption requires CRM quality or financial review, turn auto-apply off where available and record the manual decision owner. Record what happens to both arms when the experiment ends.
The launch record should capture the actual start, selected allocation, final budgets and any differences from the approved brief. Planned settings and implemented settings are separate evidence. A screenshot or export can document configuration, but it cannot establish that the experiment has produced a reliable result.
Monitor health without continually optimizing the test
Operational checks protect the experiment and the business. They answer whether the setup is serving, measurement is functioning and agreed limits are being respected. They are separate from deciding which arm is effective.
Define guardrails so that another person can apply them. “Watch lead quality” is vague. A usable rule names the qualification definition, the records to review, the review schedule, the unacceptable condition and the action owner. A spend rule needs the relevant campaigns, currency and escalation point.
Distinguish immediate faults from noisy performance. A broken destination or confirmed duplicate recording can require prompt intervention. A single expensive lead may be ordinary variation. The plan should state which conditions require immediate suspension, which require investigation and which remain observations until the scheduled review.
Consider a fictional case in which one arm’s qualified leads suddenly disappear from the report. The first action is to investigate delivery and measurement. The team should not immediately declare the treatment ineffective if its CRM import stopped. Preserve the incident period, diagnosis and restoration evidence.
If a material issue occurs, decide whether the comparison remains interpretable. Repairing both arms, restarting a test or ending it as invalid may be appropriate depending on the incident. Do not silently delete inconvenient dates. Any exclusion should have a documented reason tied to validity, with its effect on the conclusion disclosed.
Keep necessary changes in a deviation log: time, owner, reason, affected arm, expected consequence and decision about continuing. A test with documented limitations can still teach something. A test repeatedly adjusted without records leaves the reviewer unable to distinguish the proposed treatment from the operator’s reactions.
Plan for learning and delayed outcomes
The day an ad interaction occurs may differ from the day a person enquires, qualifies or purchases. Offline imports can add further delay. Recent cost may therefore be visible before the corresponding outcomes, making a partially reported period look worse than it will after it matures.
Separate campaign learning from conversion lag. Learning concerns how the tested system adapts. Lag concerns when outcomes become observable. Waiting through one does not automatically resolve the other. Record the treatment of both in the brief, using the guidance for the particular experiment type.
For bidding tests, Google recommends a ramp-up period of two weeks or three conversion cycles, whichever is longer, and excludes that ramp-up from performance evaluation. It then recommends at least 30 uninterrupted days and excluding recent days for which fewer than 90% of conversions have been reported. These are Smart Bidding test recommendations, not a universal experiment timeline or proof of sufficient evidence. Use the conversion-delay evidence for the measured outcome to set the cutoff.
In a fictional business where qualification regularly takes several days, a Friday report cannot treat every enquiry from Thursday as a completed qualification opportunity. Record the observation cutoff and how pending outcomes are handled. Apply the same rule to both arms.
Decide whether the final review will wait for the planned cohort to mature or use a consistently defined mature period. Avoid mixing dates of ad interaction, conversion and CRM import without explaining the basis. Google’s automated bidding evaluation guidance illustrates why recent periods affected by conversion delay should be treated carefully.
Save the analysis dates and report extraction time. If later conversions change the result, a reviewer should be able to explain the update instead of assuming one report was wrong.
Set a stopping rule before checking for a winner
A stopping rule states when the team will end exposure, when it will assess effectiveness and what will happen if the answer is unclear. It prevents a test from becoming an open-ended search for a favourable result.
Separate the operational stop from the decision review. A predefined safety or financial limit can stop activity immediately. The effectiveness decision should use the planned evaluation window, mature outcomes and the agreed evidence standard. Stopping for a broken form does not prove that a marketing hypothesis failed.
Write a rule such as: “At the scheduled review, adopt only if the primary outcome meets the commercial criterion, uncertainty is acceptable under the planned analysis and guardrails pass. Otherwise retain the control or record the test as inconclusive.” Include who makes that judgment.
A fixed schedule does not force a conclusive answer. If the information remains weak, “inconclusive within this budget and period” is a legitimate result. Explain whether another test is worthwhile and what would make it more informative.
If an extension is contemplated, define the circumstances and maximum exposure beforehand where possible. Do not repeatedly extend simply because the current result misses the desired threshold. An unplanned redesign or extension should preserve the original decision record and disclose the change in analysis plan.
Google’s broad match documentation describes optional automatic application of favourable results. Review that setting before launch. A platform’s automated criterion may not include your external quality, capacity or margin checks.
Read the interval, the metric and the commercial threshold together
There is no universal number of conversions that makes every experiment trustworthy. A platform threshold for eligibility or displaying results is not proof that the test can distinguish a worthwhile effect from ordinary variation.
Read the estimated difference together with its uncertainty interval. An interval spanning both useful improvement and unacceptable deterioration leaves an important decision unresolved. A result that is not statistically significant does not demonstrate equivalence. It may simply be too imprecise to distinguish the alternatives.
Reporting confidence settings also differ. Google’s general monitoring guide describes selectable levels and an 80% default. Its Demand Gen guide describes a 70% default and other options. Record the actual selected level before launch rather than assuming every report uses the same convention.
Do not lower the threshold after seeing an inconvenient result. Equally, a high selected confidence level is not a promise of commercial value. It concerns the statistical interpretation of the measured difference, subject to the design and measurement assumptions. It does not establish that the change is profitable, scalable or transferable to another market.
For applicable experiments, Google describes a bucketed statistical methodology. Advertisers cannot reproduce that method from aggregate totals alone. Adding results from separate tests and recalculating a simple percentage does not produce a valid combined significance assessment.
For a separately analysed CRM outcome, use an analysis appropriate to its assignment and data structure. Simple arithmetic can describe observed cost per qualified lead, but it does not supply uncertainty on its own. Do not present an invented interval or borrow the platform’s confidence label for a different measure.
Finally, keep practical significance visible. If an estimated saving is too small to cover the cost of implementation, the business may reasonably retain the control. That is an economic decision, not a claim that the measured difference does not exist.
Worked outcome decisions: three ways to interpret a report
All figures and intervals in this section are fictional teaching examples. They are not calculated from account data. Assume the intervals come from an appropriate planned analysis at its preselected confidence level; no confidence level is recommended by these examples. Changes are relative percentages, not percentage-point changes.
| Fictional observation | Decision and qualification |
|---|---|
| Cost per qualified lead falls an estimated 12%, with an interval from an 18% reduction to a 5% reduction. The fictional plan required a reduction exceeding 4%, and volume and quality guardrails pass. | The interval supports the prewritten commercial criterion in this illustration. A controlled adoption is reasonable after checking measurement and implementation conditions. The 4% threshold is a fictional business choice, not a general benchmark. |
| Cost per qualified lead falls an estimated 4%, with an interval from a 13% reduction to a 7% increase. | The result is inconclusive for a decision requiring a dependable reduction. The range includes deterioration. Retain the control under the stated rule and assess whether a better-powered future test is affordable. |
| Cost per qualified lead improves, but qualified-lead volume falls below the predefined operating floor. | The volume guardrail fails. Do not adopt automatically. A narrower efficiency result does not satisfy a plan that required the business to preserve useful demand. |
Consider one more fictional arithmetic example. A control spends CAD 6,000 and records 30 qualified enquiries, giving an observed cost of CAD 200 each. A treatment spends CAD 6,600 and records 36, giving about CAD 183.33 each. The observed cost reduction is about 8.3%, while qualified volume rises 20% and spend rises 10%.
Those totals describe the example; they do not establish statistical significance. Allocation, uncertainty, outcome maturity and guardrails still matter. If the plan concerned growth within an acquisition-cost limit, the additional spend may be acceptable. If the business lacked authority for that exposure, the same numbers reveal an operating failure alongside the performance observation.
Write the conclusion in the language of the tested decision. “This setting met our criterion in this campaign and period” is more defensible than “This strategy always works.” Preserve unresolved risks instead of hiding them behind a winner label.
Treat adoption as another controlled change
Before applying a result, review what the platform action will do. In a campaign replacement experiment, adoption may change which campaign receives traffic. In an asset experiment, it may change the retained assets. Google’s Performance Max experiment guide explains that application behaviour depends on the experiment type.
Save the final report, settings comparison, guardrail review and decision. Record the implementation owner, expected resulting configuration and recovery action if the change does not behave as intended. Check the resulting campaign after the authorized application rather than assuming the button implemented every intended detail.
Avoid combining adoption with an unrelated increase in budget, offer change and website redesign. Each additional change makes it harder to understand subsequent performance. If a budget change is necessary, document it as a separate business decision with its own rationale.
Monitor the adopted configuration for tracking health and commercial guardrails. A successful experiment does not guarantee the same result at a different scale or during a different season. Follow-up monitoring checks continued suitability; it does not retroactively change the original experiment’s result.
Keep a concise learning record even when the treatment loses. Rejected hypotheses, feasibility limits and measurement problems can prevent the next operator from repeating the same uninformative test.
Keep campaign improvement separate from broader causal claims
Our marketing attribution models guide explains how recorded credit differs from causal impact. Keep those reporting questions separate from the experiment decision.
A campaign experiment can provide evidence about the effect of a tested change within its assignment method, population and measurement setup. That does not automatically establish how many total sales the business gained compared with running no advertising.
Google’s Conversion Lift uses a different design to estimate conversions caused by advertising, including a group withheld from the studied ads. Eligibility and study requirements apply. Even that result must be interpreted within the outcomes, people or regions, time period and assumptions actually measured.
Similarly, a platform-reported conversion-value increase is not automatically additional business profit. It may omit costs, returns or outcomes outside the measurement definition. Use accurate language about what the experiment measured and leave broader business incrementality to a study designed for that question.
A campaign test involving destination URLs also does not settle whether the website can support a broader website experiment. Website assignment, traffic feasibility and implementation are separate design questions. Our conversion rate optimization guide provides context for investigating visitor friction without making every campaign test a website-testing programme.
Questions to resolve before launch
Can we compare this month with last month instead?
A before-and-after comparison can identify a change worth investigating. It also includes everything else that changed between periods, such as demand, competition, offers and sales capacity. It does not provide the same controlled comparison as the proposed experiment. Label it as an observational comparison and limit the conclusion accordingly.
Can we test several ideas at once?
Some platform workflows support multiple arms or bundled changes. The practical question is whether the design can answer the decision with the available information. More arms divide the available exposure. A bundle evaluates the bundle. For a first operational test, a focused question is usually easier to implement and interpret.
What if the platform metric improves but sales disagrees?
Return to the written outcome definition and evidence. Sales may have identified a real quality problem, or the two teams may be counting different things. Review the records consistently and document any measurement defect. Do not resolve the disagreement by choosing whichever dataset supports the preferred campaign.
Should an inconclusive result trigger a larger budget?
Only after a new commercial decision. More information may help, but additional spending is not automatically justified. Consider the likely value of resolving the uncertainty, the cause of weak evidence and whether a revised campaign or question would be more informative. Preserve the inconclusive result even if another test follows.
What should we bring to an experiment planning review?
Bring the proposed change, primary outcome definition, historical volume and lag, campaign eligibility notes, spending limit and decision owner. Discuss a Google Ads experiment plan with Canada Create to define what can be tested and measured before launch.
