---
title: "Understand Market Data methodology"
description: "Pave's Market Data benchmarks are built from compensation data collected directly from human resources platforms through automated, persistent connections. This article explains how that data is collected, processed, and published."
canonical_url: "https://support.pave.com/articles/understand-market-data-methodology-gv8BPtVoj4"
md_url: "https://support.pave.com/articles/understand-market-data-methodology-gv8BPtVoj4.md"
---
# Understand Market Data methodology

Pave's Market Data benchmarks are built from compensation data collected directly from human resources platforms through automated, persistent connections. This article explains how that data is collected, processed, and published.

## How data is collected

Pave connects to three types of HR platforms to gather compensation data:

* **HRIS (human resources information systems)**: Employee records including job title, level, location, base salary, bonus, and variable pay
* **ATS (applicant tracking systems)**: Candidate and offer data
* **EMS (equity management systems)**: Equity grant details, vesting schedules, and current equity holdings

These connections are automated and persistent. After a company completes the initial integration (typically about 15 minutes), data flows to Pave continuously without any manual submissions or annual survey inputs.

For a full list of supported platforms, see the Integrations section in Pave.

## How data is processed

Once Pave receives data from a connected platform, it goes through several steps before entering the compensation database:


1. **De-identification**: All data is aggregated and de-identified before it enters the compensation database. No individual employee data or company-specific compensation data is visible in Market Data. Pave's Master Subscription Agreement guarantees this protection.
2. **Job matching**: A machine learning algorithm matches each employee record to Pave's job architecture system (job families, levels, and tracks). For details on how job matching works, see [Understand job matching](understand-job-matching.md). For details on Pave's job catalog structure, see [Explore the Pave job catalog](explore-the-pave-job-catalog.md).
3. **Aggregation**: Matched records are grouped by role, level, location, and company characteristics to produce benchmark distributions.
4. **Benchmark generation**: For combinations with sufficient raw data, Pave computes benchmarks directly from the aggregated records. Where raw data is sparse, Pave's models generate [Calculated Benchmarks](understand-calculated-benchmarks.md) to fill gaps.
5. **Consistency Labels**: Every benchmark receives a [Consistency Label](understand-consistency-labels.md) based on the variability of the underlying data, giving you context on how reliably it represents the market.

## How benchmarks are published

Pave collects data from connected platforms daily, but publishes updated benchmarks monthly. New data releases are typically scheduled for the first Monday of each month that is not a US holiday.

The monthly cadence balances data freshness with quality and consistency. It also supports antitrust compliance by preventing changes from being traced to any individual company's data.

## What data is collected

**From employees and candidates** (via connected HR platforms):

* Location
* Job title and level
* Manager relationship
* Base salary
* Bonus and variable compensation
* Equity grants and holdings

**From organizations**:

* Company name and headquarters
* Industry
* Headcount
* For private companies: funding sources, valuation, capital raised, historical share prices
* For public companies: market capitalization, ticker, listing exchange, revenue

## Dataset composition

Pave's database includes thousands of companies across a range of industries, stages, and sizes. The dataset is strongest in technology and technology-adjacent sectors (fintech, medtech, etc.) with growing participation from financial services, healthcare, life sciences, manufacturing, and other industries.

For current participant details, see [Pave's Market Data participants](https://pave.com/products/market-data-participants).

## Data privacy and compliance

Pave is designed to protect the confidentiality of participating companies and their employees:

* All compensation data is aggregated and de-identified before display. No individual employee records or company-specific pay data is accessible through Market Data
* The largest single participant represents less than 2% of the overall dataset
* Monthly publication prevents tracing benchmark changes to individual companies
* No prospective or forward-looking wage data is included
* Customers cannot access competitor-specific compensation information

Pave holds SOC 1 Type 2, SOC 2 Type 2, and ISO/IEC 27001:2022 certifications, and is CCPA and GDPR compliant. Data is encrypted in transit (TLS 1.2+) and at rest (AES 256-bit). For more on Pave's security practices, see the Security section in Pave.
