Home/Solutions/Alternative Data

Data for alternative data

Alternative data with a paper trail.

Public-record and public-web panels built for research: every observation carries the date we saw it and the date the source published it, so a backtest sees what the market saw. Delivered as history plus a live feed, to S3, Snowflake or Databricks.

Most alternative data is exhaust: app panels, card aggregates, device pings, each with a provenance story that is hard to defend to a compliance officer and a panel that changes shape without warning. Ours is different in kind. It comes from county recorders and assessors, city permit offices, state and provincial licensing regulators, and public listing and professional web surfaces. Every source is nameable, every record is dated, and the methodology is written down.

Point-in-time integrity. We keep the snapshot, not just the current state. License files are retained as dated monthly snapshots. Listings are captured daily and the observation date travels with the row. Deeds and permits carry both the event date and the recording or issue date the source published, so you can model the publication lag rather than pretend it away. We do not restate history in place.

Entity resolution is the work. A regulator lists eXp as EXP REALTY LLC, EXP REALTY, LLC, eXp Realty of California, Inc. and EXP Realty LLC across four jurisdictions. Rolling those to one parent, and that parent to a ticker, is where the signal is made or lost. We maintain the rollups, expose the raw strings underneath them so you can audit every mapping, and version the mapping so a backtest is reproducible.

Compliance posture. Public record and licensed commercial sources only. No material non-public information, no consumer app exhaust, no scraped private surfaces. Personal fields can be dropped or aggregated before delivery for institutional use, and we will sign the diligence questionnaire and walk your data-sourcing team through each pipeline.

Engagements usually start with a free historical slice on one signal so your team can test it against a known series before any contract. Ask for the backfill depth on the specific signal you care about; it varies by source and we quote it honestly rather than averaging it.

What a row looks like

Agent headcount by brokerage, from the regulators, monthly. Raw office strings shown; the parent rollup is a separate versioned column.

office_string_as_filedparent_rolluplisted_entitylicenseessnapshot
EXP REALTY LLCeXp RealtyeXp World Holdings27,15608/2026
eXp Realty of California, Inc.eXp RealtyeXp World Holdings4,96708/2026
REAL BROKER LLCReal BrokerageThe Real Brokerage8,17508/2026
COMPASSCompassCompass Inc.5,24708/2026
COLDWELL BANKER RESIDENTIAL REAL ESTATE LLCColdwell BankerAnywhere Real Estate6,75108/2026

Signals researchers use

SignalData behind it
Brokerage agent headcount
Licensee counts per brokerage, monthly, straight from the regulators
Real Estate License Data
Homebuilder and contractor activity
Permit volume and value by builder, trade and metro
Building Permit & Contractor Data
Lender origination share
Recorded mortgage volume by lender, county and month
Mortgage & HELOC Data
Housing turnover and pricing
Transaction counts, median price, cash share, by geography
Real Estate Transaction Data
Rent index and concessions
Asking rent, days on market and price changes across ~970k active rentals
For-Rent Listing Data
Inventory and price cuts
New listings, months of supply, cut frequency, days on market
On-Market Listing Signals
Company headcount and hiring velocity
Workforce size and growth per company, computed weekly from 830M profiles
Company Workforce Data + Job Change Signals
Vehicle registration mix
Registrations by make, model and model year across 127.8M vehicles
Vehicle Ownership & VIN Data

From the ledger

Every licensed real estate agent in the US and Canada, refreshed monthly
62 regulators, validated live each month, matched to email and mobile
2.5M records
Residential roofers ranked by permit volume, city by city
24.2M permits clustered into a 464k-company master across five sources
464k companies

How a project runs

  1. Send the briefAudience, geography, must-have fields, refresh cadence and how you want to receive it. A paragraph is enough. We reply within one business day, usually with questions and a rough count.
  2. Get a sample and a match reportFor appends we run a slice of your own file and report the match rate field by field. For lists we send sample rows for your geography. Both are free.
  3. Agree terms and take the first deliveryPer matched record, flat license or subscription, sized to the pull. The first file arrives with a data dictionary and fill rates.
  4. Refresh on a scheduleMost programs become a standing monthly or weekly delivery. Deltas are flagged so you only process what changed.

Compliance

All panels are built from nameable public sources. Documentation lists each source, its publication cadence, our capture cadence and known gaps, so a compliance review has a paper trail.

Questions we get

What does point-in-time mean here?

Every observation carries two dates: when the source published it and when we captured it. Snapshots are retained, so a backtest can reconstruct exactly what was knowable on any past date.

How far back does history go?

It varies by panel. License snapshots and permit history run several years; listing captures run from when daily capture began. We state the history depth per panel before you license it.

How is the data delivered to a research team?

History as Parquet to S3, Snowflake or Databricks, then a live feed on the same schema. Entity mappings to tickers are provided where the underlying entity is public.

Is this compliant with MNPI rules?

Every source is a public record or a public web surface, named in the documentation. There is no insider, panel or exhaust data in these feeds.