Training and evaluation data with provenance.
Models are only as defensible as the data behind them. We package our datasets, or build new ones, with field-level documentation, sampling notes, PII handling and training-license terms.
We supply structured records for valuation and forecasting models, document text for extraction and classification, and grounded question-answer pairs for real estate and professional assistants. Where PII must be removed, we remove or synthesize it while preserving distributions. Where labels are needed, we produce them from ground truth in our own files rather than by annotation guesswork.
Evaluation sets are a specialty: held-out transactions, labeled deed text and license records that let you measure a model against reality.
How it runs
| 1 | Define Task, modalities, volume, PII rules. |
| 2 | Assemble Drawn from production files with documented sampling. |
| 3 | Document Data dictionary, provenance, license. |
| 4 | Deliver Parquet or JSONL by S3 with checksums. |
Questions we get
Is model training allowed?
Yes, under a training license. Redistribution of raw records is not included.