Playbook · GTM Data Model

The GTM Data Model playbook.

The data model your customer's AI agents run on — the semantic layer, the context graph, and the handful of definitions everything else quietly assumes. Blueprint → Build → Enable → Maintain. Start with the audio brief, then open the repo.

4Artifacts, in order
~30Definitions, owned
10 wksKickoff → two live agents
L0→L4Maturity we move them
Watch or listen first · 6 min

The audio brief.

One narrator, six minutes, the whole shape of the project — now with the deck that runs alongside it. Play it before you open anything else; it's the fastest way to get into the right mindset for a GTM Data Model build. In a hurry? Change the speed under the deck — it sticks next time.

0:00 / 0:00
Speed
1 / 1
Chapter 1 Narrated brief · generated with ElevenLabs · Download ↓
Chapters
The whole idea

When a customer says their AI isn't working, it is almost never the model. You can rent a frontier model for twenty dollars a month. What you cannot rent is a business in a form that model can read. This project is how a business becomes readable.

The Standard

Four artifacts, in this order.

Not two layers. Four — and the order is the whole method. Most failed versions of this project skipped number two and shipped a glossary.

The GTM Data Model
Build order · never parallel
1Definitions→ 2 · Compiler→ 3 · Graph→ 4 · Retrieval

Ship 1 and 2 without 3 and you have numbers you can trust with no context around them — useful. Ship 3 without 1 and 2 and you have a graph that returns confident answers about the wrong account. Ship 4 first and you have bought a very expensive search engine.

1
The definitions — meaning, agreed and owned
One file per metric. Source tool, object, field, grain, time semantics, segmentation rule, and a plain paragraph of what it means. Lives in Git. This is the artifact we hand them: their GTM Brain.
2
The compiler — definitions made executable
A semantic layer that compiles those definitions into correct SQL at query time, so every consumer resolves the same number from one place. Lives on the warehouse. Without this, the repo is documentation and every agent still guesses.
3
The graph — entities resolved and connected
Accounts, people, deals, activities and the edges between them — entity-resolved, time-aware, multiplayer. Lives in a database, never on a laptop. This is what turns a number into a story.
4
Retrieval — unstructured artifacts, anchored
Calls, emails, tickets and docs chunked, embedded, and anchored to a graph node. A property of the graph, not a fourth system. Optional, and always last.

So what

The single most common way this project fails is stopping after artifact one. A repo full of beautiful definitions that nothing compiles is a glossary, and glossaries decay in ninety days. The repo is the contract; the semantic layer is the enforcement.

Where They Are

Halfway buys nothing.

Read the customer's level in week one. Most sit at Level 1, buy tooling that requires Level 3, and conclude that AI doesn't work.

LevelNameWhat's trueCeiling — the best agent you can run
L0ReportedNumbers live in decks and spreadsheets. Two people produce two answers.A writing assistant.
L1CentralizedData is in one warehouse. Nothing is defined. Nothing is resolved.Text-to-SQL. Confidently wrong.
L2ResolvedEntities are deduplicated and joined. One account is one account.Counts you can trust.
L3DefinedEvery metric compiles from one owned definition. the barAgents you'd trust with a number.
L4RetrievableUnstructured evidence is anchored to resolved entities.Agents that cite, rank and act.

L3 is the goal of this engagement. L4 is the stretch, and only after L3 holds.

The Shape

Four phases, one motion.

Start Here

Kick off a project in one paste.

Copy this prompt, drop in the customer name, and send it to your Claude.

paste into Claude
# GTM Data Model — kick off
Clone the LeanScale GTM Data Model template repo for my customer [Customer Name]
as [customer]-gtm-brain.

Then read AGENTS.md and run the Blueprint checklist. Start with the
question inventory — do NOT start by cataloguing metrics.

Before the kickoff call, tell me:
  1. What we already know about their stack (CRM, MAP, warehouse, BI)
  2. Which GTM motions they appear to run, and the evidence
  3. The 20 questions I should walk into the room proposing
  4. Which of the three build paths I should be steering toward, and why
1
Phase 1 · Weeks 1–4

Blueprint.

Half of this engagement is alignment, and all of that alignment happens here. Blueprint is not a discovery doc — it is a set of decisions with names attached to them. If you leave this phase without a named human owning every definition, the rest of the project is decoration.

1
The question inventory — before anything else
Twenty questions the customer wants an agent to answer, in their words, gathered from the exec team. Then define only the metrics those twenty questions need. Opening with "let's define every GTM metric" is how this becomes a ten-week project with two hundred definitions and zero working agents.
2
The motion map
Enterprise, velocity, product-led, partner. Each motion gets its own funnel, its own entry event and its own definition of qualified. Most customers believe they run one motion. They almost never do.
3
The roll-up legality matrix
For every metric × every motion: does this aggregate, aggregate with a caveat, or become a lie? This is the two-hour argument that ends the two-year argument.
4
~30 definitions, one owner each
Sales, marketing, customer, company. Full record for each — see the anatomy below.
5
The identity spine design
The match ladder for people and companies, the survivorship rules, and the review-queue thresholds. Designed here, built in phase two.
6
The maturity read and the build path
Score them L0–L4, pick one of the three paths, and get the fork signed before a line of anything is built.

Anatomy of a definition

Nine fields. A definition missing any one of them will be re-litigated within a quarter.

FieldWhat goes in itWhy it's there
NameQualified OpportunityThe word the business already uses. Don't invent vocabulary.
OwnerOne named human, customer-sideNot a team. Not us. See the rule below.
ServesWhich of the 20 questionsA definition that serves no question shouldn't exist yet.
SourceSalesforceSystem of record. Be honest when the real answer is a spreadsheet.
ObjectOpportunityWhere it physically lives.
Field / expressionStageName entered DiscoveryDown to the field. "Stage is past discovery" is not a definition.
GrainOne row per opportunityThe fan-out bug that makes every dashboard double-count.
Time semanticsPeriod (entered-in-window), not as-ofThe single most common silent error in GTM reporting.
Segmentation ruleNever roll up across motionsStraight out of the legality matrix.

The ~30 that matter

The starting set. Add only what a question demands.

Sales
8 definitions
  • Qualified opportunity
  • Stage entry & exit
  • Sales cycle time
  • Win rate
  • Reopened opportunity
  • Pipeline coverage
  • Territory
  • Quota & attainment
Marketing
6 definitions
  • Lead
  • M.Q.L.
  • Source
  • Channel
  • Campaign
  • Attribution model
Customer
7 definitions
  • Customer
  • Live date
  • Renewal
  • Expansion
  • Churn
  • Health score
  • Support escalation
Company
7 definitions
  • Account
  • Person
  • Segment
  • Bookings
  • ARR / MRR
  • Fiscal calendar
  • Employee & rep

The roll-up legality matrix

Fill this in the room, with the exec team watching. The verdict column is what you're actually there for.

MetricEnterpriseVelocityProduct-ledVerdict
ARRContract valueContract valueContract valueAlways rolls up
Win rateOpp-based, SAL→WonOpp-based, MQL→WonSignup→PaidNever — different denominators
Sales cycle timeFirst meeting → WonMQL → Wonn/a — no cycleNever — different start events
Pipeline coverage3× forward quarter3× forward quartern/aWithin motion only
M.Q.L.Hand-raise + fitScore thresholdActivation eventNever — count, don't sum
CACLoaded, incl. SE timeLoadedPaid + product costBlended only with the split shown
Net revenue retentionCohort by live dateCohort by live dateCohort by live dateAlways rolls up

The rule that holds it all up

Every definition has exactly one named human owner — a person, not a team, not a committee — and that person works for the customer, not for us. We own the process; they own the meaning. We facilitate the fight between the CRO and the CMO about what counts as qualified. We never settle it, and we never quietly settle it by writing our own definition when nobody is looking. A data model whose definitions belong to LeanScale dies the week we roll off.

The build path fork

Decide this in Blueprint, in writing. Changing paths in week six costs a month.

Path A · Vasco
Fastest to L4
  • Choose when: sub-$50M, no warehouse, no data team
  • Semantic layer comes from marking GTM objects — motions, channels, functions, ICPs, dimensions
  • Graph, entity resolution and provenance are built in
  • Agent surface ships with it (MCP)
  • Trade: their model shape, their opinions
Path B · Assemble
Fully theirs
  • Choose when: they have a warehouse and an analytics engineer
  • dbt Semantic Layer or Cube on Snowflake / BigQuery / Databricks
  • Graph-in-warehouse, native vector search
  • No lock-in, no vendor definitions
  • Trade: slower, and they must staff it
Path C · Map what exists
Most common at scale
  • Choose when: they already own Looker, a CDP, or a mature dbt project
  • LookML is a semantic layer — don't rebuild it
  • Map what's there onto the four artifacts, fill only the gaps
  • The gap is almost always the graph and the agent surface
  • Trade: you inherit their debt
2
Phase 2 · Weeks 3–9

Build.

One design principle governs everything in this phase: the repo is the source, the warehouse is the runtime. Definitions are authored, reviewed and versioned in Git, then compiled into the semantic layer by CI. Nobody edits a metric in a BI tool ever again.

The GTM Brain

What we hand them. One clone and their Claude has the whole model.

[customer]-gtm-brain/
# the contract — human-agreed, machine-readable
definitions/       one file per metric; owner in the frontmatter
questions/         the 20 questions, each linked to the definitions it needs
motions/           motion map + the roll-up legality matrix

# the model — what CI compiles and deploys
models/            transforms (dbt or equivalent)
semantic/          the semantic layer spec — the compiler's input
graph/             node + edge schema, identity ladder, survivorship rules
retrieval/         chunking, anchoring and embedding config  (optional, last)

# the surface — how humans and agents reach it
agents/            the shipped agents, read-only by default
CLAUDE.md          how to reason about this business
.mcp.json          the endpoints an agent may call
docs/              workshop outputs, decisions, the arbitration log

The compiler

Tool-agnostic. Pick for what they already run, not for what's fashionable.

Semantic layer engines
  • Cube — best when agents are the primary consumer; REST/SQL/GraphQL APIs, open-source core
  • dbt Semantic Layer — best when they already run dbt; definitions sit beside the transforms
  • Looker / LookML — already a semantic layer if they own it; awkward to expose to agents
  • Vasco — the layer and the graph together, on Path A
The non-negotiable
  • Never let a model author business logic.
  • Text-to-SQL is the anti-pattern — the model silently guesses joins, grain and definitions, and is wrong in a way nobody can see
  • Text-to-semantic-query is the pattern — the agent picks from an approved menu, the semantic layer compiles it correctly
  • Policy lives at the definition, so access control can't be bypassed by asking differently

Identity resolution — the hard part

Not the metrics. Deciding when two records are the same thing. Build it as a ladder, never as a guess.

Person ladder
  • 1 · Verified email, exact auto-merge
  • 2 · Normalized email — strip +tags, gmail dots auto-merge
  • 3 · LinkedIn URL / provider person ID auto-merge
  • 4 · Phone, E.164 review queue
  • 5 · Normalized name + resolved account review queue
Account ladder
  • 1 · Root registrable domain auto-merge
  • 2 · Provider IDs — SFDC 18-char, HubSpot company ID auto-merge
  • 3 · Registry ID — DUNS, PDL, Clearbit auto-merge
  • 4 · Normalized name + country review queue
  • + Explicit parent/child edges for subsidiaries and acquisitions

Do this or lose the database

Block the free-mail and shared domains before you turn resolution on. Without a blocklist, every @gmail.com lead merges into one enormous fake account, and you will find out in front of the CRO.

Rule 1

A permanent surrogate key

Every resolved entity gets a ULID that is stable forever, never reused, and never derived from a business attribute. Emails change. Domains change. Companies rebrand. The key must not.

Rule 2

Every merge is reversible

The resolved entity is a new node with edges back to each source record — never a destructive overwrite. Get this wrong and a bad merge rule corrupts a customer's data with no path back. This is the one that ends engagements.

Rule 3

Every edge carries time

Valid-from and valid-to, not just created-at. Without it, "who owned this deal last quarter" returns today's owner, confidently. This is the number one reason a graph gives the wrong answer.

Rule 4

Provenance on everything

Source system, source record ID, confidence, as-of — on every node and every edge. No provenance means no citation, which means no trust, which means no adoption.

Where the graph lives

It must be multiplayer — hosted, authenticated through their IdP, with one governed write path. Never a local file.

Default · ~80% of customers
Graph-in-the-warehouse
Resolved entity tables plus an edge table — from_id, to_id, edge_type, valid_from, valid_to, source, confidence — in the same warehouse as the semantic layer. Traverse with recursive CTEs or BigQuery's native graph syntax. Multiplayer by default, same auth, same governance, one less system. GTM graphs are small: hundreds of thousands of nodes, not billions.
Only if
A dedicated graph database
Neo4j AuraDB or Neptune — justified when routine queries are three or more hops: org-chart traversal, multithreading path-finding, partner referral chains. Costs you another system, another sync and another access model. Running Neo4j for 200k accounts is maintenance with no benefit.

Anchored retrieval

✕
You never vectorize the database
Structured records are answered by the semantic layer, exactly and correctly. Embedding a table of ARR values so you can ask "what's our ARR" is the signature AI-project mistake.
✓
You vectorize unstructured artifacts only
Transcripts, emails, tickets, notes, docs. Every chunk carries its metadata and is anchored to a graph node.
✓
Retrieval is graph-scoped first, vector-ranked second
Filter to this account's last ninety days, then rank by similarity inside that set. Pure vector search across the corpus is what returns a beautifully relevant paragraph about the wrong customer.
✕
Vectorization is not an agent
Chunk, embed, upsert, index — on a schedule. That's ETL. The one place an agent belongs is typed entity extraction: pulling pain, risk, next step, champion and economic buyer out of a transcript, each with a verbatim quote as evidence. Deterministic work doesn't get an agent; judgment does.
Vector store — pick the boring one
  • Whatever is native to their warehouse — Snowflake Cortex Search, BigQuery vector search, or pgvector on Postgres
  • One less system, and the graph-scope filter is a plain SQL WHERE — which is exactly what anchoring needs
  • Pinecone / Weaviate / Qdrant only when scale or latency genuinely demands a dedicated store

Write-back governance

Default posture
  • Every agent ships read-only
  • Writes go through a scoped action layer, never direct CRM API access
  • One named human owns each write scope
  • Every write is logged with the agent, the prompt and the resulting diff
Proof dashboards
  • Definition coverage — how many of the 20 questions are answerable
  • Resolution match rate, by ladder rung
  • Review-queue depth and age
  • Model integrity score — unowned definitions, orphaned fields, stale sources

Week six · steal this

Ask the executive sponsor for one number out of the semantic layer that contradicts their current board deck. If nothing contradicts, nothing has actually been defined — you've rebuilt their old reporting with better tooling. Something always contradicts. Find it early and the whole engagement earns its trust in a single meeting.

3
Phase 3 · Weeks 8–10

Enable.

We lead this project, which means Enable is handover, not training wheels. The customer supplies two roles we cannot fill for them — the executive sponsor who arbitrates definitions, and the CRM admin who holds the org. Everything else transfers.

The discipline

Ship two agents, not ten. The first agent carries all the infrastructure; by the tenth it's an afternoon of work. But hand them ten on day one and nobody adopts any of them. Pick the two that answer questions the sponsor actually asks on a Monday morning, prove those, and let demand pull the rest.

Week 8
Owner training
Every definition owner walks their own file. Changing a definition now means a pull request with their name on the approval.
Week 9
Agents 1 & 2 live
Read-only, cited, answering two of the twenty questions. Demoed by the sponsor, not by us.
Week 10
Office hours + handover
Repo cloned into their Claude. Arbitration path documented. Keys transferred.
Week 14
30-day review
Which definitions changed, who approved them, and what the review queue looks like.
4
Phase 4

Maintain.

Data models don't break loudly. They drift. So Maintain is a watchdog, not a calendar — these are the triggers that should page someone.

Trigger · people

An owner leaves

Their definitions become orphans the same day. An unowned definition is a dead definition — reassign within a week or it will be quietly contradicted by a spreadsheet.

Trigger · resolution

Match rates move

A ladder rung whose match rate shifts more than a few points means an upstream format changed, or a new source started writing dirty records. Catch it before the review queue floods.

Trigger · corporate

They acquire a company

The single most reliable way to break entity resolution. New domains, duplicate accounts, a second CRM, and suddenly the parent/child edges matter more than anything else in the graph.

Trigger · usage

A metric stops being queried

The most useful signal on this list. It almost always means somebody rebuilt it in a spreadsheet because they didn't trust yours. Go find out why — that's a definition problem wearing a tooling costume.

Trigger · upstream

Schema drift

A Salesforce field renamed, a picklist value added, a HubSpot property deprecated. CI should fail the build the moment a definition points at something that no longer exists.

Trigger · GTM

They launch a new motion

Every new motion invalidates part of the roll-up legality matrix. Reopen it, don't patch around it — a motion added without a matrix update is how "win rate" starts lying again.

On the clock
  • Quarterly — re-resolve entities in full and diff against the prior run
  • Quarterly — definition review with owners; retire what no question needs
  • Monthly — integrity score in front of the sponsor
  • Continuous — CI on every definition change, with the owner as the required approver