Data for ai & machine learning
Real-world records for models that have to be right.
Valuation models, real estate assistants, document extraction and forecasting engines need training data with provenance. We package ours with documentation, PII handling and license terms written for model training.
We build from production files rather than scraped text, which means labels come from ground truth: the sale price on the deed, the license status from the regulator, the permit type from the city. Evaluation sets are held out by time so a model is scored against the future, not the past.
Plays that work
| Play | Data behind it |
|---|---|
| AVM training Transactions, characteristics, permits and listings by parcel | Real Estate Transactions + Property Records + Building Permits |
| Document extraction Deed and mortgage text with structured labels | AI Training Data: Real Estate |
| Assistant grounding Question-answer pairs grounded in license and property records | AI Training Data + Real Estate License Data |
| Entity resolution Labeled person and company match pairs | Identity Resolution + Global Professional Profiles |
How a project runs
- Send the briefAudience, geography, must-have fields, refresh cadence and how you want to receive it. A paragraph is enough. We reply within one business day, usually with questions and a rough count.
- Get a sample and a match reportFor appends we run a slice of your own file and report the match rate field by field. For lists we send sample rows for your geography. Both are free.
- Agree terms and take the first deliveryPer matched record, flat license or subscription, sized to the pull. The first file arrives with a data dictionary and fill rates.
- Refresh on a scheduleMost programs become a standing monthly or weekly delivery. Deltas are flagged so you only process what changed.
Compliance
Training sets ship with a data card: sources, dates, PII treatment, known gaps and license terms. Where a source's terms prohibit model training we exclude it and say so.
Questions we get
What makes a training set defensible?
Provenance. Every record names its source and carries the date we observed it, and the license terms state that training is a permitted use. Scraped text cannot offer either.
How are evaluation sets built?
Held out by time. A valuation model is scored on sales recorded after its training cutoff, so the test is against the future rather than a random slice of the past.
How is personal data handled?
Person-level fields are removed, hashed or replaced with synthetic values depending on the task, and the treatment is documented in the data card. Property and company records are not personal data in most jurisdictions.
Can you produce labeled documents?
Yes. Deeds, permits and license records paired with their structured fields make ground-truth extraction sets. Volume and geography are scoped per brief.