PostHog Handbook Library / Growth

1,645 words. Estimated reading time: 8 min.

Data foundations

Auto TL;DR

At a Glance

This long page covers these main areas. The list is generated from the article headings, so it updates with every handbook rebuild.

  1. Why do companies need data tools?
  2. The modern data stack, at a glance
  3. So what's wrong with normal?
  4. Traditionally, a different vendor for every step
  5. PostHog collapses the stack
  6. How PostHog maps to the stack
  7. How users grow into the data stack
  8. What's special about a context warehouse?

Full deck for sales training [Recorded sales training session]

Why do companies need data tools?

Every single SaaS platform or tool you use generates data. Payments in Stripe, customers in the CRM, and app data in a production database. That data could be used to ask questions that will help you improve your product, learn more about customers.

Some questions customers might have:

| Question | Data it needs | |----------|---------------| | Which features drive revenue? | PostHog Events + Stripe | | Which accounts are ready to upsell? | PostHog Usage + Stripe plans | | What do customers do right before they cancel? | PostHog Events + Stripe/Chargebee | | Do support tickets predict churn? | PostHog Usage + Zendesk/Intercom | | Which leads deserve the sales team's time? | PostHog Usage + Hubspot/Salesforce | | Which channels bring customers that stick? | PostHog Events + Stripe + Meta Ads |

The modern data stack, at a glance

So many tools to ask a simple question:

  1. Sources – where data is born
  2. Ingestion – move it in to
  3. Storage – the warehouse
  4. Modeling – clean & shape it so it's accurate
  5. Orchestration – keep it running
  6. BI – analyze it
  7. Activation – push it back out

So what's wrong with normal?

A traditional stack is six or more separate tools, each with its own bill, login, and pipelines between them all to maintain.

Product data usually lives in a completely separate world from business data.

Traditionally, a different vendor for every step

| Step | Traditional tool | |------|------------------| | 1. Sources | Sources are your own, no one to replace. | | 2. Ingestion | Fivetran, Airbyte, Portable | | 3. Storage | Snowflake, Redshift, BigQuery, Databricks | | 4. Modeling | dbt, sqlmesh | | 5. Orchestration | Airflow, Dagster | | 6. BI | Looker, Hex, Tableau, Metabase | | 7. Activation | Hightouch Segment, RudderStack|

Product analytics is usually your biggest data source so most of your data is already in PostHog, if customers use our full stack you don't have to export that data anywhere.

PostHog collapses the stack

Instead of assembling multiple tools, PostHog brings the layers into one platform.

The context warehouse = your events and your business data, together and queryable under one roof.

How PostHog maps to the stack

| Step | PostHog | |------|---------| | Sources | Autocapture + SDKs | | Ingestion | Warehouse Sources | | Storage | Managed Warehouse | | Modeling | Data modeling + SQL editor | | Orchestration | Scheduled syncs & materialization | | Analysis & BI | Notebooks, Insights & dashboards | | Activation | CDP, batch exports & reverse ETL |

How users grow into the data stack

Users only need data tools once they start generating a reasonable volume of data. They will adopt PostHog's other products first.

  1. Start · pre-data – PostHog core products: Analytics, session replay, feature flags, experiments
  2. First data tool – Connect Stripe: Revenue data, joined to product events. Having Stripe synced signals that a user has customers!
  3. Step 2 – Ask PostHog AI: Questions about their data, joining together PostHog events and external data
  4. Step 3 – Build a dashboard: Revenue + product in one saved view
  5. Step 4 – Provision a warehouse: A user reaches a volume of warehouse to need a dedicated warehouse (10k+ Event rows)
  6. Long term – Model & scale: Once you have a warehouse, it needs modeling to keep the data usable and build better dashboard

This runs from an early-stage start-up through to a more mature company with a first data hire.

What's special about a context warehouse?

Definition: A context warehouse is data storage and tooling, optimized for agents as the consumer.

What a context warehouse is not:

How to talk about context warehouse vs data stack

Companies talk about a "data stack" like a "tech stack." We sell the whole stack, all-in-one but you can bring your own tools if you like, and we're optimized for agents.

| The traditional data stack – humans interpret & act | The context warehouse – agents interpret & act | |-----------------------------------------------------|------------------------------------------------| | Assembled from many separate tools | One all-in-one stack, bring your own tools if you want | | You build and maintain the pipelines between them | Ingestion, storage, modeling, and intelligence | | Built for humans to query, dashboard by dashboard | Agents get the context to know what a customer or feature flag is | | People read the data, then decide what to do | Query in natural language or using SQL | | Optimized for analysts | Optimized for agents as the consumer |

Who the context warehouse is for

It's our ICP, but more data focused.

Primary persona · today: Product engineers

Engineering-led startups that want to run their own data before they hire a data team.

Seed–Series B · 15–500 people · no data hire yet

Cares about

Frustrated by

Secondary persona · next: Data Lead

The first data hire at a growing PostHog customer will usually be a generalist who can build infrastructure and query it.

Series A–C · solo data team · already on PostHog

Cares about

Frustrated by

Not yet: data scientists. They're a company's second data hire. Most of our customers don't have data scientists, and our modeling tools aren't mature yet. Once we have a product that can compete with dbt, we can start targeting data scientists.

Why are we targeting these groups

Product engineers: We want to educate this user so that they adopt our products and build data foundations, including setting up warehouse sources, before hiring a dedicated data person…

Data leads: …so that when this person gets hired, the stack that already exists is PostHog, and they will give us a shot instead of churning to tools that are more established

How to talk to them

Start with treating them like a human being.

Product engineers

Primary message

"Ship features faster and understand users deeply without needing to hire a data person straight away"

Why they'll care

Lead with

Data leads

Primary message

"Stop maintaining the stack. Start doing the work."

Why they'll care

Lead with

Acronyms, decoded

An appendix of sorts.

| Acronym | Meaning | What it is | |---------|---------|------------| | ETL | Extract, Transform, Load | Pull data out from the source, clean it, then load it into your warehouse. | | ELT | Extract, Load, Transform | Load raw to the warehouse first, clean it once it's there. The modern default. | | CDC | Change Data Capture | Live updates only the data that is new or changed to your warehouse | | OLTP | Online Transaction Processing | The database that runs your app. Fast, one row at a time. | | OLAP | Online Analytical Processing | The database built for analytics. Slower, millions of rows at once. | | SQL | Structured Query Language | The coding language you use to ask a database questions. | | BI | Business Intelligence | Dashboards and reports built on top of your data. | | CDP | Customer Data Platform | Unifies customer data and pushes it to your other tools. | | Reverse ETL | – | Send modeled warehouse data back out into apps so it's up to date and enriched in tools like Salesforce. | | DAG | Directed Acyclic Graph | A map of the order modeling tasks need to run in |

Canonical URL: https://posthog.com/handbook/growth/sales/data-foundations

GitHub source: contents/handbook/growth/sales/data-foundations.md

Content hash: 6bca8dd4c562d2bd