A GA4 audit should end with a clear answer to a business question: which numbers can we use for this decision, under what conditions, and what must wait? A property can collect events and still give an unreliable answer about qualified enquiries, returning customers or a change in campaign performance. Conversely, two reports can disagree for legitimate reasons without either implementation being broken.
The useful output is a set of evidence-backed decisions, not a percentage that labels the entire property healthy. This guide covers definitions, reporting consistency, freshness, identity, thresholds, sampling and consent-modelled data. It provides an evidence checklist and a severity rubric you can adapt to your reporting responsibilities.
Example and evidence note: every business, count and audit finding in the worked examples is fictional. The method draws on public Google documentation checked on 8 October 2026. It does not describe an audit of a real GA4 property. The companion examples have synthetic consistency checks; they do not demonstrate that a website, consent platform or live integration works.
Start with the decision, not the dashboard
Write down the action someone wants to take before opening reports. “Understand performance” is too broad to audit. “Decide whether September’s completed online enquiries justify reviewing the October media allocation” gives the work a population, period, outcome and consequence. It also exposes whether GA4 alone can answer the question.
A useful decision statement names the business outcome, the exact metric proposed as evidence, the comparison being made, the decision owner and the date the decision is needed. Include what an error could change. A number used to choose a page for further research needs a different level of assurance from a number used to calculate a supplier’s performance payment.
For a fictional Ontario equipment supplier, a completed quote-request event might support investigation of website demand. It cannot, by itself, establish that sales accepted those requests, that the enquiries came from serviceable regions, or that any order was profitable. Those are different outcomes with different evidence requirements. Keep the distinction visible even if the dashboard uses the convenient label “leads”.
Define the smallest scope that supports the decision. That might be one production website, one completed calendar month and three approved enquiry events. Record excluded apps, other brands, test environments and offline enquiries. A property-wide audit should examine common settings and reporting dependencies, but each conclusion must still name the population it covers.
Agree acceptable uncertainty before seeing the result. Avoid a universal “within five per cent is fine” rule. A small discrepancy can reverse a close campaign ranking; a larger explained difference between two unlike populations may be harmless. Ask whether the uncertainty could change the action, and document the reasoning. If nobody can explain the uncertainty, classify the number as not assessed or hold the affected use.
Build an evidence checklist someone else can repeat
Begin with a dated snapshot of the current setup and the relevant history. Use only access authorised for the audit. Request redacted evidence when screenshots or exports would otherwise expose personal information. A screenshot of a total is weak evidence unless another reviewer can reconstruct its dates, filters, dimensions and metric definition.
| Audit area | Evidence to retain | Question it resolves |
|---|---|---|
| Property and coverage | Property identifier, Standard or 360 tier, production streams, included domains or apps, property time zone and reporting currency | Does the population match the business question? |
| Definitions | Exact metric and dimension names, outcome definition, scope, key-event counting method and effective dates | Does this number mean what the reader thinks it means? |
| Collection assurance | Dated results from authorised representative journey tests, with expected and observed outcomes | Is there evidence that the required journey is observable? |
| Reporting conditions | Date range, extraction time, report surface, filters, segments, comparisons and quality messages | Can someone reproduce the result and recognise limitations? |
| Identity and consent | Reporting identity, modelling status, applicable dates and consent implementation evidence | Which users and behaviours can the report represent? |
| History and availability | Retention setting, relevant changes, filter state, collection gaps and known incidents | Are periods comparable and still available? |
| Decision control | Finding, severity, permitted use, restriction, owner and reopening condition | What should the business do with this evidence? |
We plan and run B2B marketing across search, ads, content and email for Canadian companies.
Record an evidence location, not a claim that a check “looks good”. A suitable reference can be a restricted screenshot name, an approved test record or a saved report specification. Keep sensitive evidence in the organisation’s controlled storage; the public example register uses fictional references only. Do not paste customer emails, full form payloads or visitor identifiers into a shared audit sheet.
Separate observed facts from explanations. “The report has a thresholding warning” is an observation. “The small audience probably caused it” is a hypothesis until supported. “Do not rank those audience rows” is a decision. Combining all three into one confident sentence makes an uncertain diagnosis look like established fact.
Also record what was unavailable. If a reviewer cannot see a setting or obtain a prior export, that gap is part of the audit. An empty evidence cell should never quietly become a pass.
Audit metric meanings before investigating discrepancies
Write a short metric dictionary for the numbers that drive decisions. Include a business definition, the GA4 field used, its unit, the population, the relevant time basis and an explicit exclusion. “Completed web enquiries, measured by the approved confirmation event, excluding known test traffic” is more useful than “conversions”. The dictionary should also state which parts still require sales-system verification.
Google distinguishes total, active, new and returning users. For example, total users and active users are different metrics with different qualification rules; a dashboard’s short “Users” label is not enough to establish equivalence. Check the selected field rather than assuming that every report counts the same population. See Google’s user-metric definitions.
Keep counts and rates distinct. A count of completed enquiry events is not a count of people who enquired. A rate needs a named numerator and denominator. If you divide enquiries by sessions in an external worksheet, label that calculation exactly. Do not present it as a native GA4 key-event rate unless it follows that metric’s definition. Someone can trigger multiple events within one session, so apparently similar calculations can answer different questions.
Inspect key-event counting methods and their history. Google documents counting once per event and once per session; changing the method applies to future key events rather than rewriting past data. A period spanning a change therefore deserves an explicit comparability review. Record the method for each material key event rather than assuming the property has one universal rule. Defaults can depend on how the key event was created: migrated Universal Analytics goals can use once per session, while other key events default to once per event. Check the actual setting and its effective date. Google’s counting-method guidance explains the distinction.
Also distinguish the Event count metric from Key events. A key-event counting choice applies to the latter; it does not redefine a raw event-count query. Record whether the outcome was marked as a key event during the comparison period. A custom metric label in a worksheet should still identify which source metric and event selection produced it.
Do not use a reporting change to conceal a collection problem. If the defined outcome is one accepted submission and a controlled test records two, the audit should flag the affected outcome and request a focused diagnosis. This guide stops at documenting the risk, affected dates and evidence needed for release. Detailed duplicate-event repairs, cross-domain changes and tag deployment procedures belong in their own implementation work.
For monetary values, record the exact revenue metric and currency. Separate a value assigned to an enquiry from money collected from a customer. A hypothetical CAD 150 value attached to a lead event is an assumption, not proof of CAD 150 revenue. Mixed definitions can make a tidy chart commercially unusable even when every event arrives.
Check scope and comparison conditions
Traffic-source dimensions have different scopes. “First user” dimensions describe initial acquisition; “Session” dimensions describe acquisition for sessions; event-scoped source dimensions support attribution of key-event credit. Google explains these distinctions in its traffic-source scope documentation. Preserve the full field name in the audit, including its prefix.
Imagine a fictional visitor first finding a supplier through organic search and returning later through an email campaign. A first-acquisition report and a session-acquisition report can classify that activity differently because they ask different questions. Their disagreement is not sufficient evidence of a tracking fault. Before proposing a correction, describe the question each report answers.
Compare one meaningful pair of reports at a time. Align the property, streams, dates, time zone, metric, dimensions, filters, comparison groups and relevant attribution settings. Save the specifications together. For timestamp-based exports, document a start-inclusive, end-exclusive interval in the named property time zone. For September in the fictional Toronto example, that is 1 September at 00:00 up to, but not including, 1 October at 00:00. Retain the time zone as well as the extraction offset; daylight-saving changes can make fixed-offset comparisons misleading. If one report includes all events while another selects a particular key event, fix the comparison specification before interpreting the difference.
Google also documents expected differences between reports and explorations, including filtering, supported fields, date ranges, modelling and processing. Two screens with similar titles are not necessarily equivalent queries. Use its reports-versus-explorations explanation when classifying a discrepancy.
Review boundaries that can change independently of marketing performance: a second website entering the property, an app launch, an enquiry-definition revision, a new filter or a different reporting identity in saved extracts. Build a short timeline. If the definition changed halfway through September, a single month-over-month percentage may blend two measurement regimes.
Where feasible, compare stable periods on either side using the same definition. Where that is impossible, label the series break and avoid a precise improvement claim. Do not manufacture a historical adjustment factor from one convenient week. A documented interruption in comparability is a useful audit finding, even when there is no recoverable historical answer.
Treat freshness as part of the number
“September enquiries: 84” is incomplete without an extraction date and reporting context. The same query can change as processing finishes. Google says processing can take 24–48 hours, reports and explorations may not be synchronised, and attribution credit for key events can change for up to 12 days. Its typical availability times are not guarantees. See data freshness and processing guidance.
Use separate reporting rules for operational monitoring and settled comparisons. An operations team may accept a provisional view to investigate a possible outage. A monthly channel-allocation review needs a documented cut-off and a policy for restatements. Choosing a later cut-off is a business reporting convention; it does not make every field permanently final.
In a fictional check, the team extracts the same completed-month enquiry count twice and sees 80 followed by 84. The second is four higher, or five per cent above the first. That arithmetic proves only that the saved outputs differ. It does not prove a collection fault, identify which records changed or establish that the later figure will never change again.
Record both values, both extraction times and the identical query conditions. Classify the earlier report as provisional if the evidence supports that explanation. If the later change could reverse a planned action, hold that action until the owner has reviewed the cause and agreed a suitable reporting basis.
A good recheck instruction names a time, a saved query and a decision rule: repeat the completed-period query at the agreed review point, compare material movements, and record whether any unresolved change affects the intended use. “Wait a bit” leaves the same argument for the next meeting.
Establish what a user count represents
GA4’s current reporting identity choices are Blended, Observed and Device based. Blended uses User-ID, device ID and modelling in that order; Observed uses User-ID and device ID; Device based uses device ID. Google says this reporting choice does not change collection or processing and can be switched without permanently affecting the underlying data. See reporting identity.
For the audit, retain the identity setting attached to the report or extract. Do not describe a device-based user total as a verified count of individual customers. Similarly, a reported user total is not automatically a count of CRM contacts, households or purchasing companies. Those populations require their own definitions and evidence.
Ask the implementation owner whether User-ID is used, which authorised business state supplies it and what checks support consistent assignment. This is a request for evidence, not permission to export identifiers or inspect customer records. A generic value shared by many visitors would be a different risk from normal differences between anonymous devices and signed-in accounts.
If two saved reports use different identity settings, disclose that before describing a movement as growth or decline. A controlled comparison may help explain sensitivity, but document its purpose and preserve the original reporting basis. Choosing the option that produces the most flattering number is not a quality improvement.
Decide what the number can support. A consistently defined user series may help identify a period worth investigating. It may still be unsuitable for an exact customer-reach promise or person-level reconciliation. State that boundary on the decision record so the caveat travels with the exported number.
Keep zero, missing and suppressed values separate
A displayed zero is not sufficient evidence of zero activity. Verify that it represents a measured result for the requested scope, rather than a spreadsheet fill, unavailable history or an export convention. A blank cell may mean that the query was not run, the result was unavailable or a value was withheld. These are different states and should remain different in every export and calculation. Use a separate value-status field; leave the numeric field empty when no usable number exists. In these templates, observed means a numeric output has been verified with its availability context. It is not Google’s distinction between observed and modelled data. Record sampling or modelling separately in the quality note.
Google applies data thresholds to reduce the risk of identifying individuals or sensitive information from reports. Thresholds are system-defined, and the data-quality indicator identifies affected reports or explorations. A wider period may sometimes provide sufficient aggregation, but that also changes the question being answered. Follow Google’s thresholding guidance; do not infer or reconstruct withheld individual values.
For a fictional regional audience comparison, suppose one segment displays a number while another is withheld. You cannot conclude that the withheld segment produced no enquiries. Nor can you fairly rank the two by replacing the withheld value with zero. Mark the second value as suppressed and hold the ranking. A broader, appropriately aggregated analysis might support a different decision after its own checks.
Keep GA4’s special dimension labels intact as well. Google describes (not set) as a missing dimension value and (data not available) as a distinct condition that can involve processing or unavailable traffic-source information. They are not synonyms for direct traffic or zero activity. See its explanation of unavailable data.
Record the affected dimension and query. “Campaign information unavailable” is a narrower conclusion than “all data is wrong”. A completed enquiry total may remain useful while its campaign allocation is on hold. The allowed use should be no broader than the evidence supports.
When a denominator is missing, suppress the derived rate too. When a denominator is actually zero, the ratio is undefined; label it accordingly. Neither case justifies displaying a zero per cent rate. Otherwise the spreadsheet creates apparent certainty that the source never supplied.
Distinguish sampling from grouped rows
Sampling estimates a result from a subset of data. Google says sampling can affect reports, explorations or requests when relevant limits are exceeded, and the data-quality icon shows the proportion used. Unsampled results can still use approximation for distinct user and session counts. “Unsampled” therefore does not mean every metric is an exact person-by-person ledger. See Google’s sampling documentation.
As checked on 8 October 2026, Google documents an event-level query limit of 10 million events for Standard properties. Analytics 360 starts at 100 million per query and supports up to 1 billion with the more-detailed Explore option. These are query limits, not collection caps or guaranteed precision. Recheck the current documentation and property tier before relying on them. Record the message and sample proportion for the exact query, not a remembered property limit. A sampled result may help identify a broad pattern for follow-up. It needs more scrutiny if the proposed action depends on a small difference between two low-volume groups. There is no universal sample percentage in this guide that certifies a decision.
As an audit procedure, simplify an over-detailed question only if the simpler question still serves the business need. Removing dimensions or selecting a shorter period can produce a different result with different limitations. Preserve both specifications and check their indicators. Do not combine overlapping extracts or add daily user counts to imitate a deduplicated monthly total.
The (other) row is a different mechanism. Google groups less common dimension values when a supporting table exceeds its row limit. The warning can remain relevant even when a filter hides the visible row. Google’s explanation of the other row connects this behaviour to cardinality and query complexity.
Consider a fictional report that groups many page variants into “other”. The overall figure might answer a high-level volume question after other checks. It does not identify which individual low-volume page should be removed. Record the granularity limit and request a suitable page-level evidence source before making a page-level decision.
Do not deduct points for these conditions merely because an icon looks cautionary. Evaluate whether the condition obscures the comparison that matters. A report can have a valid aggregate use and an invalid detailed use at the same time. Your decision register should allow both conclusions without forcing a property-wide pass or fail.
Separate observed and consent-modelled evidence
Consent changes what can be observed. Google’s behavioural modelling estimates activity using patterns from users who consent to identifiers. Eligibility depends on implementation and sufficient suitable data; meeting the published prerequisites does not guarantee eligibility. Inspect the current modelling status and its effective date rather than assuming that a consent banner or Blended identity guarantees modelled results. See behavioural modelling for consent mode.
Write down whether the selected output includes estimates, excludes them or has unavailable estimated data. Unsupported views and segments can differ from an aggregate report. Do not interpret a modelled total as a recoverable list of people or enquiries. It cannot identify which particular visitor would have appeared in a CRM system.
Google distinguishes basic and advanced consent-mode implementations. Basic mode prevents Google tags loading before consent-banner interaction and keeps them blocked when consent is denied; advanced mode can send cookieless pings under denied consent. These implementations create different observable inputs. Use the current consent-mode documentation to identify the implemented approach and request evidence from the responsible owner.
The audit should not recommend weakening a consent choice to make a chart look more complete. Record technical behaviour separately from the organisation’s assessment of its privacy obligations. This guide does not establish that any particular implementation satisfies Canadian law or that additional collection is appropriate.
For a fictional month containing a consent-platform change, split the review around the known change date. Ask whether observed collection, the reporting population or modelling availability changed. If those effects are unresolved, a before-and-after user comparison should carry a caveat or be held. Calling the difference “lost demand” would jump beyond the evidence.
If a report includes estimates and a separate export does not, do not force the totals to match by adding a guessed uplift. Retain each output’s meaning. Modelling can support aggregate analysis while remaining unsuitable for verification of an individual lead or transaction.
Review history, retention and excluded data
Some audit questions cannot be answered retrospectively. Google’s retention setting affects user-level and event-level availability for explorations and funnel reports, not standard aggregated reports. Increasing retention cannot restore data already deleted. User-level and key-event retention options are 2 or 14 months; other event data can have 26, 38 or 50 months in 360. Age, gender and interest data retain a two-month limit. Property-size limits can also shorten event-level retention. Do not assume that a tier’s longest option applies to every dataset. Check the current setting and the dates available for the specific analysis, using Google’s retention documentation.
A separate availability rule matters for Google Ads data in Analytics: Google documents a rolling 36-month limit, with older Ads values shown as zero in daily, hourly or weekly line charts. Full-calendar-month monthly charts have different treatment. Check the chart and date range before classifying an older zero; preserve unavailable history as missing in the audit register.
Do not promise a year-over-year journey analysis simply because a standard report displays an older monthly total. Record the earliest usable date for the required surface and fields. If historical detail is unavailable, narrow the question, use an already authorised independent record or mark the analysis unavailable. A blank historical period is not evidence of no activity.
Review active collection filters and their effective dates. Google warns that excluded data from an active internal-traffic data filter is not processed and will not later be available in Analytics or BigQuery. A report filter is a different operation. Its internal-traffic filtering guidance explains this distinction and the testing state.
This makes the evidence timeline essential. A drop following filter activation may reflect changed coverage. A before-and-after total needs that context even if the filter was correctly configured. Record the intended exclusion, who owns it and the evidence showing what it covered; do not change an active filter as a casual audit experiment.
Finally, inventory known outages and omitted journeys. An apparently stable property total can hide a missing mobile flow or a recently introduced form. Ask for dated test evidence for each important journey and classify untested coverage explicitly. A successful desktop test is not a universal certificate for the property.
Reconcile differences without pretending systems are identical
Choose a reference that answers the same business question wherever possible. A CRM can help verify accepted enquiries; a commerce system can help verify orders. Neither automatically matches GA4’s population, dates or definitions. Before comparing, write a short reconciliation contract covering the outcome, date basis, inclusion rules, permitted identifiers, time zone and known coverage gaps.
Use an aggregate comparison first. In a fictional supplier example, an approved operations report contains 96 completed web enquiries for a month, while the saved GA4 output reports 84 approved enquiry events. The numerical difference is 12, and 84 divided by 96 is 87.5 per cent. Call that a comparison ratio. It is not automatically a tracking accuracy rate.
The totals alone do not establish that 12 enquiries were lost by GA4. The systems may treat consent, repeat enquiries, rejected submissions or time boundaries differently. They may also contain genuine errors. A responsible conclusion distinguishes “difference observed” from “cause established”. Hold completeness claims until comparable evidence resolves the relevant uncertainty.
Where record-level reconciliation is authorised, keep identifying records in the approved system and minimise what leaves it. The audit register can reference a controlled reconciliation record without copying personal data. Do not require new identifiers or new data transfers merely to finish the worksheet.
BigQuery is also a separate reporting surface, with different processing and modelling characteristics. Google provides a comparison of Analytics reporting surfaces. BigQuery does not include GA4 behavioural modelling, key-event modelling or data-driven attribution. Its export is therefore not a report-equivalent ledger. A calculation should document its logic, export coverage and available fields; the absence of GA4 query sampling in BigQuery does not establish complete collection.
If the two sources cannot be aligned, keep them side by side with their definitions. The honest deliverable may be “GA4 supports directional website engagement analysis; the operations report remains the reference for accepted enquiry counts.” That is useful guidance without inventing a bridge between unlike datasets.
The broader commercial interpretation belongs in a separate discussion. Canada Create’s PPC ROI guide covers advertising returns. Establish measurement suitability here before using those numbers to judge commercial performance.
Use severity and decision status for different jobs
Severity describes the consequence of a finding. Decision status describes whether a particular use is supported. Keep them separate. An unresolved regional dimension may be a serious blocker for regional allocation while leaving an independently checked property total useful for another question.
| Severity | Meaning | Response |
|---|---|---|
| Critical | Evidence indicates an immediate data-handling concern or a materially invalid number already driving a consequential action. | Escalate to the accountable owner immediately and contain the affected use through the approved process. |
| High | An unresolved problem could reverse or materially alter the pending decision. | Hold that use and specify the evidence required to reconsider it. |
| Moderate | A bounded limitation affects interpretation, but a narrower supported use remains. | State the restriction and obtain the decision owner’s acceptance. |
| Low | A documentation or maintenance gap has no demonstrated material effect on the current decision. | Assign an owner and review date without overstating the impact. |
| Unrated | Evidence is insufficient to judge the consequence. | Gather evidence; do not interpret the absence of a rating as low risk. |
| None | No unresolved issue was found within the documented checks for this use. | Retain the scope and reopening conditions. |
Apply one of four decision statuses. Use means evidence supports the named action within its stated scope. Use with caveats means a restricted interpretation is supportable and its limits must travel with the number. Hold means evidence shows a material unresolved limitation for that action. Not assessed means the necessary evaluation has not been completed. None of these statuses means “GA4 is universally correct”.
Do not average severity labels into a score out of 100. Several low-impact documentation gaps cannot cancel one material defect. Likewise, count of issues is not impact: one issue may affect every executive report, while ten others concern unused fields. Report the decisions affected and the action needed.
A complete finding has six parts: the observation, its evidence reference, the affected use and period, the severity rationale, the current restriction and the release condition. “Fix tracking” fails this test. “Hold the September regional ranking until an appropriately aggregated report supports a comparable view of both regions” gives the owner a specific path forward.
Keep the decision owner and repair owner distinct where appropriate. An analyst can explain a limitation; a business owner accepts the remaining uncertainty; an implementation owner handles a scoped correction. Closure requires evidence that the release condition was met, followed by a fresh decision. Completing a task is not enough to certify the number.
Work through a fictional audit decision
Consider Cedar Example Supply, a fictional Canadian business with one production website. Its owner wants to review next month’s media allocation. The audit covers September 2026 completed online quote requests. Calls, offline enquiries, sales acceptance and revenue are excluded. Every record described here is invented for instruction.
The reviewer first separates three questions: whether the completed enquiry total is suitable for a settled monthly report, whether regional groups can be ranked, and whether GA4 captured every operational enquiry. These are separate decisions even though the original dashboard presented them together.
- Monthly volume: two fictional saved outputs contain 80 and 84 events. The later export is documented with the agreed definition and reporting conditions. Until the extraction difference is reviewed, the earlier total remains provisional. A subsequent decision to use the later total must name its scope and must not imply complete coverage of all enquiries.
- Regional ranking: one required value is suppressed. The numeric field stays blank, its status is “suppressed”, and the ranking is on hold. The release condition is suitable comparable aggregate evidence, not a guess about the hidden row.
- Completeness: a separate fictional operations report contains 96 enquiries. The 12-count gap is observable, but the reason is unresolved. The reviewer holds any “we track all enquiries” claim and assigns the definitions-and-coverage comparison to the measurement owner.
The owner can still organise further investigation without reallocating the budget based on the incomplete ranking. No blanket statement that the property is broken is necessary. Equally, an apparently plausible overall total cannot approve the unsupported regional or completeness claims.
Write the resulting decision in ordinary language: “Use the documented September GA4 total only for the reviewed website-event trend. Do not equate it with accepted sales enquiries, infer suppressed regional activity or claim full coverage against operations records. Reopen the decision if the saved query changes, the definition changes or reconciliation reveals a material error.”
That statement is the practical product of the audit. It tells a colleague what they may do today, what they must avoid and what evidence would change the answer. A long list of settings without this conclusion leaves the actual decision to guesswork.
Use the evidence register and decision templates
Download the GA4 data quality audit kit (ZIP) for the blank templates, fictional examples and source register described below.
The companion set contains a blank CSV evidence register, a fictional CSV example, a blank JSON decision card, its fictional counterpart, a field dictionary, instructions and a synthetic validation record. The formats are intentionally simple. They are documents for reviewing evidence, not connections to Google Analytics and not imports into a live property.
Start with one decision card. Enter the question, population, period, owner and reporting conditions. Then add one evidence row for each check. Use the same decision identifier across rows supporting that decision, but a unique check identifier for every row. If one saved output informs two decisions, create a separate decision-specific row that cites the same evidence and identifies the reuse. It is not another independent observation or another count to add. This makes it possible to see which findings block a decision without combining unrelated outcomes.
A compact copyable outline for an evidence record is:
Decision:
Metric and full dimension names:
Population and reporting period:
Report surface and extraction time:
Value status: observed / missing / suppressed / not_assessed
Numeric value: leave empty unless observed
Quality message and evidence reference:
Severity and rationale:
Decision: use / use_with_caveats / hold / not_assessed
Permitted use and restriction:
Owner and release condition:In the CSV, preserve an empty numeric cell for missing or suppressed results. In JSON, preserve null. Enter an actual numeric 0 only after checking that the source reports an available zero for that scope. A filled, suppressed or unavailable zero must not be promoted to an observed count. Treat status labels as part of the evidence, not cosmetic formatting. Spreadsheet import settings should preserve identifiers as text and dates in their documented format.
The fictional files demonstrate a changing count, a suppressed value, a missing value, an observed zero, a sampled result and an unassessed check. Their validation covers structure, allowed statuses and the illustrated arithmetic. It does not verify a live tag, consent state, identity configuration or Google report. A real audit needs separately authorised evidence for those checks.
Before sharing a completed copy, review the evidence references and notes for confidential information. Keep a public training copy fictional. Retain the completed organisation-specific copy wherever your team controls access, retention and revisions.
Keep the decision current after the audit
An audit conclusion has a shelf life determined by its dependencies. Reopen relevant decisions after changes to event definitions, consent behaviour, streams, identity, filters, reporting fields or connected business processes. Also reopen them after an unexplained material movement or when the business asks a different question of the same number.
Keep a short decision history rather than overwriting the old conclusion. Preserve the prior status, new evidence, effective date and reason for change. If a report already circulated with an unsuitable figure, identify which conclusions need correction. Reissuing a total without explaining its effect can leave the original decision in place.
Choose a review cadence based on how often the implementation and decisions change. A launch period may need frequent review; a stable monthly report may need a recurring check plus change-triggered reassessment. This is an operating choice, not a universal platform requirement.
Keep neighbouring audits focused. Canada Create’s technical SEO audit guide addresses crawlability, indexing and related website checks. Those findings may explain site performance, but they do not certify GA4 definitions or reporting suitability.
Common questions about GA4 audits
Can we use GA4 when some checks fail?
Yes, if evidence supports a narrower use and the failed checks do not undermine it. Approve the specific decision, metric, population and period. A caveat that changes the meaning of the result belongs beside the number, not hidden in an appendix.
Does disagreement between reports prove an error?
No. First compare the definitions and conditions, then examine quality messages and processing. If the difference remains unexplained and could change the action, hold that use. Avoid declaring either report correct solely because it matches expectations.
What should the completed audit deliver?
A reproducible evidence register, a definitions and change record, a prioritised set of findings and explicit decisions about permitted use. Each unresolved finding needs an owner and a release condition. The audit should reduce uncertainty about action even when some historical questions remain unanswered.
