How to Build a First-Party Identity Graph in BigQuery for Marketing Attribution

If you run paid media for an ecommerce brand on Shopify, you've probably had this conversation: Meta says it drove the sales. Google says it drove the sales. GA4 says 40 percent of your revenue came from “Direct” and “Unassigned,” which is another way of saying “no idea.” Finance wants to know which of the three to believe before approving next month's budget.

None of those systems is lying. They just can't see the same customer. The person who tapped your ad in Instagram, got pushed into iPhone Safari, came back two days later on their laptop, and bought after clicking an email looks like three or four different people to your analytics stack. Every one of those breaks is a place where credit falls into the “Direct” bucket or gets handed to whichever channel happened to touch the customer last.

A first-party identity graph is how you stitch that customer journey back together. This article explains what one is, what it's built from, how it works in practice on a Shopify + GA4 + BigQuery stack, what it does and doesn't fix, and how it compares to the hosted attribution tools most brands are choosing between.

What is a first-party identity graph?

A first-party identity graph is a table in your own data warehouse that links every identifier you legitimately collect about a customer (analytics cookie IDs, ad click IDs, a hashed email from a pop-up, a customer ID from checkout) into a single persistent person ID. Once that table exists, sessions, ad clicks, and orders connected by those identifiers can be joined to the same person, so attribution reflects the whole journey rather than the last fragment of it.

The important concept is first-party. The graph is built only from data your site collects directly with the customer's consent. It lives in your BigQuery project, not a vendor's. You can query it, audit it, and feed it into any reporting, ad platform, or model you like.

That is the difference between an identity graph and a “pixel.” Similar subscription-based attribution tools also resolve identity, but the identity-resolution logic stays inside their platform. 

Why GA4 attribution breaks for Shopify brands

We covered the mechanics in detail in Shopify Attribution: The 10 Ways Your Revenue Gets Misassigned. Almost every failure mode in that article is, at root, an identity problem:

  • In-app browser handoffs. A tap on a Meta or TikTok ad opens the in-app browser. The customer taps “Open in Safari” or comes back later in their normal browser. New cookie, new client_id, and the click ID never made it across.

  • Cross-device journeys. Research on the phone, buy on the laptop. GA4 sees two unrelated users unless one of them is logged in.

  • Cookie expiry and consent. Safari caps JavaScript-set cookies at seven days. Consent banners suppress tracking until accepted, if at all. Returning customers routinely arrive with a fresh identity.

  • Checkout and payment redirects. Third-party payment flows, app-based checkouts, and Shop Pay can interrupt session continuity on the way to the order.

  • Email and retargeting as last-click magnets. Because these channels reach people who already know you, they tend to be the last touch. Without a persistent ID, GA4 can't tell you that the same person came in from paid social four days earlier.

GA4's own session stitching handles some of this if you pass a user_id, but only for logged-in users, only from the point you start passing it, and only within GA4's reporting limits. The BigQuery export gives you the raw events. What it doesn't give you is the person.

The identifiers you already have

Most Shopify brands running GA4 and server-side tagging are already collecting everything a graph needs. The work is capturing it consistently and joining it. Here is what goes into a typical build, ranked by how much you can trust each signal.

Identifier Where it comes from What it links Confidence
Shopify customer ID Data layer when a customer is logged in, and at checkout Every session where the customer is recognized, plus the order Deterministic
Checkout email (hashed) Shopify checkout / order webhook The purchase to the person Deterministic
Pop-up email (hashed) Discount or newsletter sign-up form, hashed in the browser before it's sent The anonymous session that produced the sign-up to the same person's later purchase Deterministic
Google click ID (gclid) Landing page URL, persisted in the _gcl_aw cookie The ad click to the session, and to later sessions on the same browser Deterministic, single-use
Meta click ID (fbclid / _fbc) Landing page URL, persisted in the _fbc cookie The ad click to the session and later sessions Deterministic, single-use
GA4 client ID (user_pseudo_id) GA4 _ga cookie Sessions in the same browser Browser-level
Meta browser ID (_fbp) Meta pixel cookie, ideally set server-side for longevity Sessions in the same browser; also the key Meta matches on for CAPI Browser-level
Server-side fingerprint ID (e.g. Stape User ID) Server-side GTM, derived from IP, user agent, and TLS signals Sessions that look like the same device Probabilistic, low confidence

Three practical recommendations on capture:

  • Hash email in the browser, never send raw email to GA4. When the pop-up submits, take the address, lowercase and trim it, SHA-256 it, and send the hash as an event parameter (Google’s User Provided Data field / variable is for Google’s own use and hashes only for that purpose). 

  • Read the ad cookies server-side. _fbp, _fbc, and _gcl_aw aren't in the GA4 BigQuery export. A server-side GTM container (we use Stape) can read them on every request and write them to a separate identity-events table in BigQuery. That keeps your raw identifiers out of GA4's parameter limits and out of GA4's data retention rules.

  • Treat fingerprint IDs honestly. A server-side “user ID” built from IP and user-agent will merge household members and office colleagues, and it carries different consent obligations under UK GDPR and PECR than a login or an email. It's useful as a tie-breaker with a short time window. It is not a primary key.

How the graph is built in BigQuery

The build runs as a scheduled batch, usually daily, and produces one table: a mapping of every raw identifier to a single person_id. The logic is simple to describe and fiddly to get right.

Step 1: Extract identity edges. For every event in the GA4 export and the identity-events table, pull out every pair of identifiers that appeared together. A session with a user_pseudo_id, a gclid, and a hashed pop-up email produces three edges connecting those three values.

Step 2: Filter the edges. Each edge gets a confidence tier and a time window. Deterministic edges (customer ID, hashed email) are retained according to the client’s approved retention policy. Click-ID edges are kept for 30 days and dropped if the same click ID shows up against more than a handful of browsers, which usually means a shared link. Fingerprint edges get a window of a few days and a low weight.

Step 3: Find connected components. Any two identifiers connected by a chain of surviving edges belong to the same person. This is a graph traversal, and it's the reason the approach is called an identity graph. In BigQuery it's an iterative query that converges in a few passes.

Step 4: Write the person ID back. Every session, click, sign-up, and order gets a person_id column. From this point, “which channel brought this customer in?” is a GROUP BY, not a guess.

-- Simplified: first touch per person, after stitching
SELECT
  g.person_id,
  ARRAY_AGG(
    STRUCT(s.session_start, s.source, s.medium, s.campaign)
    ORDER BY s.session_start LIMIT 1
  )[OFFSET(0)] AS first_touch
FROM sessions s
JOIN identity_graph g
  ON g.identifier = s.user_pseudo_id
GROUP BY g.person_id

The output feeds three things: a rebuilt channel attribution model in BigQuery (first, last, or position-based, with the “Direct” bucket largely emptied), a new-versus-returning revenue split that actually reflects returning customers, and a clean person-level event stream for the ad platforms.

What changes once the graph exists

Attribution you can defend

The most visible change is the shrinking of “Direct” and “Unassigned.” Sessions that GA4 couldn't source are now joined to a person whose first session had a gclid or an fbclid. Credit moves from “we don't know” to a real channel.

The second change is in retargeting and email. In a working identity graph, these channels should show far fewer first-touch conversions, because a higher proportion will have a prior visit.

The third is new versus returning. Once the same person is recognized across their second and third purchases, ROAS reporting stops crediting a Meta ad with “acquiring” a customer who bought from you last quarter.

Example from a recent build

A US trade association sells event tickets through a desktop checkout. Most of its paid social traffic arrives on a phone. To GA4, the person who tapped the ad and the person who bought the ticket were two different users.

Over a five-month campaign, the graph weighed 1.75 million candidate links, resolved around 280,000 users, and joined 17,596 of them across identifiers no cookie could connect. It also rejected 1,534 potential merges where combining records would have fused two known contacts into one.

The result: 71% of purchase journeys gained sessions no single identifier could reach, and 20% of purchase journeys picked up a paid-media touchpoint that would otherwise have gone unattributed. By the final week before the event, with months of identity evidence behind it, those figures were 78% and 25%, respectively.

This suggests that, before the ID graph, most purchase journeys were incomplete in GA4, with meaningful portions of the customer path missing from view. It also shows that paid media was materially under-attributed, since many purchase journeys gained paid-media touchpoints that would otherwise have gone uncredited.

Better signals back to the ad platforms

This is the part that improves ad performance rather than just reporting on it.

When a stitched person converts, you know the gclid from their original Google click, the _fbc and _fbp from their Meta click, and their hashed email. Send the purchase to Google Ads with the click ID and hashed email via Enhanced Conversions, and to Meta with the click IDs and hashed email via the Conversions API. The platforms get conversions they had lost to the Safari handoff, their match quality improves, and their bidding algorithms have more accurate data to optimize against.

The same person table also gives you cleaner audiences: a suppression list of recent purchasers that actually contains recent purchasers, or a lookalike seed built from high-LTV people rather than high-LTV cookies.

A better foundation for incrementality and MMM

In our guide to attribution, incrementality and MMM, we argued that attribution is the micro view and the other two methods calibrate it. An identity graph doesn't replace incrementality testing. It makes the attribution inputs to those tests credible, and it gives you person-level cohorts (first purchase date, acquisition channel, LTV) that make geo tests and cohort analysis far easier to run.

How much of the problem does it actually fix?

Not all of it, and anyone who tells you otherwise is selling something.

Published benchmarks from warehouse-native attribution vendors put the stitching rate for converting customers at roughly 35 percent when the only signals are click IDs and IP, rising to around two-thirds when a login ID or hashed email is added. Our experience on Shopify builds is consistent with that range. Strict consent configurations push it lower; a well-placed email capture pushes it higher.

The right way to judge the graph is not the coverage number but the delta: how much revenue moved out of “Direct” into a real channel, and how much the platform-reported and warehouse-reported numbers converged. Those are the figures a CFO can act on.

Three ways to get an identity graph

Most Shopify brands weighing up attribution end up choosing between three approaches.

  1. A hosted attribution platform
    These add their own tracking pixel to your store and resolve identity inside their platform. Most now export data to your warehouse, so attributed orders and channel metrics can land in BigQuery for further analysis. What stays with the vendor is the resolution itself: the pixel that captures the signals, the rules that decide when two visits are the same person, and the assumptions baked into the numbers. You take on another tag on the site, another version of the truth to reconcile against GA4 and Shopify, and a subscription that usually scales with revenue or spend. For a brand that wants a working dashboard on day one and has no analytics engineering capacity, that's a fair trade.

  2. A warehouse-native identity resolution product
    Architecturally the closest to what this article describes: deterministic stitching that runs inside your BigQuery and writes to your tables. These are built and priced for brands spending well into six figures a month on media, with the onboarding and managed-service model to match.

  3. A graph built from the stack you already run
    For brands on Shopify, GA4, and BigQuery that spend enough on paid media for attribution to matter but sit below that enterprise threshold, the practical option is to build the graph from what's already there: the GA4 export, the server-side GTM container, and the checkout and pop-up data you're collecting anyway. No additional pixel. Stitching logic that's readable SQL code you can inspect and change. One set of numbers, consistent with the rest of your measurement. And no recurring license for the privilege of seeing your own data.

This is what Deducive builds. The question isn't whether you can get the data out; most tools will export something. It's whether the identity logic is yours: built from your signals, on your rules, in your project, so the numbers hold up when finance asks how they were made.

Privacy and consent

An identity graph is built from personal data and has to be treated that way.

  • Consent gates capture. Click IDs and ad cookies are advertising identifiers. Under GDPR and PECR (and equivalents elsewhere), they sit behind marketing consent. Server-side capture must respect the same consent state as the browser tags, not just read whatever cookies happen to exist.

  • Hash at the edge. On-site email capture is hashed in the browser before transmission; checkout email received server-side is hashed before being written to the identity graph. The warehouse stores hashes, not addresses.

  • Fingerprinting is a different category. IDs derived from IP and device signals are harder to justify under consent rules than a login or an email the customer typed in. We weight them accordingly and recommend clients take a view from their privacy adviser before relying on them.

  • The customer's rights still apply. Because the graph is in your warehouse, deletion and access requests can be handled properly, giving you direct control over access and deletion workflows.

What a Deducive identity graph engagement looks like

A typical build for a Shopify brand covers:

  1. Audit of existing GA4, GTM, server-side tagging, and consent setup, including a baseline measurement of Direct and Unassigned share.

  2. Capture implementation: hashed email from pop-ups and checkout, customer ID on login, click IDs and ad cookies via server-side GTM, written to BigQuery.

  3. Stitching logic in BigQuery with confidence tiers, time windows, and fan-out filters, scheduled daily.

  4. Rebuilt attribution reporting: channel attribution with a shrunk Direct bucket, new versus returning revenue, and the retargeting first-touch test as QA.

  5. Platform feedback: Enhanced Conversions and Conversions API payloads enriched from the graph.

  6. A stitching-rate report so you can see, month by month, what proportion of converters the graph is resolving and where the gaps are.

Everything is delivered into the client's own Google Cloud project.

Frequently asked questions about ID graph

What is an identity graph in marketing attribution?

An identity graph is a data structure that links all the identifiers associated with one customer (cookie IDs, click IDs, hashed email, customer ID) into a single person record, so that sessions and conversions from the same person can be attributed together rather than treated as separate users.

Do I need an identity graph if I already use GA4?

GA4 attributes at the session and cookie level. If a meaningful share of your revenue shows as Direct or Unassigned, or if cross-device and in-app-browser journeys are common for your customers, GA4 alone will misassign revenue. An identity graph built on the GA4 BigQuery export fixes the identity layer that GA4 is missing.

Can I build an identity graph without BigQuery?

You can build one in any warehouse, but for GA4 users BigQuery is the natural choice because GA4's raw event export lands there natively and Google Cloud scheduling and SQL handle the daily batch without extra infrastructure.

Is a first-party identity graph GDPR compliant?

It can be, provided capture is gated by consent, email is hashed before storage, and probabilistic signals such as fingerprinting are handled conservatively. Because the data stays in your own warehouse, access and deletion requests are easier to honor than with a hosted vendor.

How is this different from 3rd party subscription-based attribution systems?

Those platforms resolve identity with their own pixel and their own rules, and report the result in their dashboard or export it to your warehouse. A first-party graph is built from the signals you already collect, with stitching logic you can read and change, inside your BigQuery, and it produces one set of numbers consistent with GA4 and Shopify rather than a parallel one.

What stitching rate should I expect?

Typically between a third and two-thirds of converting customers, depending on how many deterministic signals you capture. Email capture and logged-in checkout are the biggest levers.

Will an identity graph improve my ROAS?

Indirectly, yes. It improves the conversion data sent to Google and Meta, which improves their matching and optimization, and it gives you attribution you can actually reallocate budget on. It won't make an unprofitable channel profitable, but it will show you which channels those are.

Measurement you own

If your Shopify brand is spending meaningful money on paid media and a large share of your revenue is showing up as Direct or Unassigned, the problem is almost certainly identity, and it is fixable inside the tools you already have.

Deducive's Marketing Attribution and Server-side Analytics services cover the identity graph build end to end, from capture to rebuilt attribution reporting, delivered into your own BigQuery.