# B2B Buyer Lab — Data Index

**Folder:** `~/Documents/B2B Buyer Lab/data/` · Curated 2026-08-31 · All JSON validated on indexing.
Every record carries source URLs and accessed dates. `"not publicly available"` is a recorded value, not a blank. All case-study numbers are provider-claimed unless a verification layer says otherwise. Provider eligibility follows the same criteria across the directory.

## Authority order (which file wins)

1. **`agencies-*.json` + `tools.json`** — the source of truth. Hand-curated evidence records; edit here.
2. **Enrichment layers** (pricing-atlas, consolidation-events, case-study-audit, interlocks, buyer-economics) — derived analyses with their own live verification; where an enrichment layer corrected a source record (see pricing-atlas `verification_log`), the correction is authoritative until synced back.
3. **Master tracker** (`master-*.{jsonl,csv,json}`) — GENERATED from layer 1 by `build-master-tracker.mjs`. Never edit by hand; regenerate: `node build-master-tracker.mjs`.
4. The site (`../site/src/data/`) holds a copy of layer 1 + tracker summary; `../site/public/data/` publishes the validated catalog and generated exports. After source edits run `(cd ../site && npm run data:refresh && npm run build)`.

## Layer 1 — Provider datasets (21 categories, 881 records, 809 unique providers — v3, 2026-08-31)

> v3 note: `index.json` carries the live per-file counts and is validator-enforced. Current totals: 21 files, 881 records (525 full + 356 stubs), 809 normalized provider IDs, and 485 public full-record profile routes. New since the original 15-file table: `agencies-intl-emea.json` (43) · `agencies-intl-canz.json` (34) · `agencies-intl-apac-latam.json` (32) · `agencies-social-influencer.json` (39) · `agencies-video-creative.json` (32) · `agencies-sales-abm-localization.json` (39); all 15 original files were also densified.

| File | Records | Category | Research doc |
|---|---|---|---|
| `agencies-geo.json` | 41 | GEO / AEO agencies (flagship) | `../research/01-…` |
| `agencies-ai-seo.json` | 25 | AI-SEO / technical search | `../research/02-…` |
| `agencies-digital-pr.json` | 20 | Digital PR (7 sell AI-citation PR) | `../research/03-…` |
| `agencies-seo.json` | 32 | B2B SEO (20 full + 12 stubs) | `../research/07-…` |
| `agencies-content.json` | 29 | Content marketing (19 + 10 stubs) | `../research/08-…` |
| `agencies-paid-media.json` | 30 | Paid media / PPC (20 + 10 stubs; fee models) | `../research/09-…` |
| `agencies-email-retention.json` | 28 | Email/retention incl. Klaviyo tiers (18 + 10) | `../research/10-…` |
| `agencies-platform-crm.json` | 33 | HubSpot + Salesforce partners (21 + 12) | `../research/11-…` |
| `agencies-platform-build.json` | 31 | Shopify + Webflow (19 + 12; 100% dir-verified) | `../research/12-…` |
| `agencies-dev-design.json` | 33 | Dev shops + design studios (21 + 12) | `../research/13-…` |
| `agencies-analytics-ops.json` | 28 | Analytics/CRO/RevOps (18 + 10) | `../research/14-…` |
| `agencies-commerce-platforms.json` | 27 | Adobe Commerce/BigCommerce/Woo (17 + 10) | `../research/15-…` |
| `agencies-enterprise-systems.json` | 26 | Dynamics/NetSuite/ServiceNow/Atlassian (16 + 10) | `../research/16-…` |
| `agencies-cloud-data.json` | 27 | AWS/GCP, Snowflake/Databricks, AI-ML (17 + 10) | `../research/17-…` |
| `agencies-cms-martech.json` | 28 | WordPress/Drupal/headless/martech (18 + 10; 94% verified) | `../research/18-…` |

Record schema: `RECORD-SCHEMA.md`. Breadth stubs are identified by one of the explicit source fields `stub`, `record_status`, `record_type`, or `record_depth`; the generator normalizes these to `record_depth: "stub"`. Across the 21 files there are 525 full provider-category records and 356 breadth stubs. Low confidence alone does not make a record a stub. Normalized-name dedup gives 809 provider IDs; only the 485 unique full-record routes are published as profiles pending alias/acquisition reconciliation (see `MASTER-TRACKER.md`).

## Layer 2 — Tools dataset

| File | Records | Contents |
|---|---|---|
| `tools.json` | 18 | AI-visibility/AEO monitoring tools: pricing tiers, engines, collection method (API/browser/hybrid/undisclosed), trials, verified affiliate terms | `../research/04-…` |

## Layer 3 — Enrichment & verification layers (2026-08-31)

| File | Inner contents | Feeds |
|---|---|---|
| `pricing-atlas.json` | 145 normalized price points · 16-category synthesis · 5 independent benchmarks · 20-entry live verification log (4 source corrections) | Pricing benchmark pages · `../research/20-…` |
| `consolidation-events.json` | 93 ownership/mortality events (2016–2026) · 9 dead-brands-in-live-listicles exhibits · 2 good-hygiene counter-examples | Consolidation Index study · `../research/21-…` |
| `case-study-audit.json` | 25 boldest claims audited: 2 corroborated / 7 checkable / 16 uncheckable; figure-drift cases | "We Audited the Claims" page · `../research/24-…` |
| `interlocks.json` | 24 verified ownership/advisory/publishing interlocks + 31-practitioner credibility table (in research doc) | Disclosure surfaces · `../research/25-…` |
| `buyer-economics.json` | 4 evidence blocks (contract norms, freelancer rates, in-house cost, tool-stack cost) + TCO table at 3 company sizes | Tool-vs-agency-vs-in-house framework · `../research/26-…` |
| `publishable-insights.json` | 3 dated aggregate claims with buyer implications + editorial/commercial boundary rules | Homepage ticker, llms.txt |

`verified-claims.json` now contains 55 adversarially checked claims from the self-listicle, tier-inflation, and dead-brand review (`../research/19-…`).

`metro-presence.json` now maps 889 presence entries across 154 metros from the original 15 provider files (`../research/22-…`). The six international/new-category files are explicitly queued for the next location pass, so this layer is useful but not yet worldwide-complete.

## Layer 4 — Generated master tracker (do not hand-edit)

| File | Contents |
|---|---|
| `master-provider-category-tracker.jsonl` | 438 provider-category records, one JSON per line, evidence-enriched |
| `master-provider-category-tracker.csv` | Same records, spreadsheet-ready |
| `master-providers.json` | 401 normalized provider IDs with profile routing + evidence-layer counts |
| `master-tracker-summary.json` | Generated totals: category coverage, verification/self-ranking/ownership-flag counts |
| `build-master-tracker.mjs` | Manifest-driven generator. Regenerate after any layer-1 edit: `node build-master-tracker.mjs` |
| `validate-data.mjs` | Strict catalog validator: parsing, manifest counts, stub counts, source-list presence and aggregate identity checks |
| `MASTER-TRACKER.md` | Field guide, guardrails, reconciliation status |

## Machine-readable manifest

`index.json` in this folder is the build manifest as well as the machine-readable catalog (file → layer → record counts → status → research doc). Validation and tracker generation read it directly; a file is not silently promoted into the trusted dataset merely because it happens to exist in this folder.

## Geographic scope

This is designed to become a worldwide provider source, but current density is strongest for agencies serving the United States and English-speaking markets. HQ and office strings are navigational signals, not proof that an agency actively serves every nearby market. Country, metro and service-market claims require explicit verification before they can support a location ranking. `metro-presence.json` will become the authority for curated location claims when its pending verification pass lands.

Global expansion order: normalize legal entities and aliases → add structured country and service-market fields → verify metro presence → publish a coverage report by country/category → open under-covered markets for researched submissions. “Worldwide” describes the intended coverage model, not a current completeness claim.

## Refresh cadence

Tools: quarterly minimum (consolidation is monthly-fast). Agency profiles: quarterly. Pricing points: re-verify before any public citation (the atlas verification log shows drift happens). Consolidation events: rolling — add on discovery. Dead-brand listicle exhibits: re-fetch before publication (pages change).
