Skip to content
01AI training data

Better training starts with better data.

Structured retail datasets and custom web collection for AI training, fine-tuning, and evaluation. Define the sources and fields you need, review a sample, and explore your delivered records in your client workspace.

02What we deliver

From your specification to usable records.

01

Custom collection to your specification

Start with the model task, then define the sources, fields, date coverage, and volume your dataset needs.

  • Source set, schema, and intended training use agreed before collection
  • Sample batch for your review before the full delivery
  • Geographic, product, and time coverage scoped to your project
  • One-time datasets or recurring deliveries prepared by our team
02

Structured retail training data

Product records, historical prices, and availability for model training, fine-tuning, and evaluation.

  • Product identifiers and available catalogue attributes
  • Dated price and availability observations at store level
  • Explore granted datasets in your client dashboard
  • Download CSV, JSON, or Parquet
03

Available source metadata

Understand where records came from and when they were observed. Missing or unchecked metadata stays explicitly unknown.

  • Source URLs and observation dates or supplied capture times
  • Separate ingestion timestamps for delivery tracking
  • Import run identifiers for tracing a delivered batch
  • Rights signals included when checked; unknown values are not permission
04

Preparation for your training workflow

Agree the quality checks and dataset preparation your project requires before ordering.

  • Review field completeness, coverage, and a representative sample
  • Discuss annotation, taxonomy, and deduplication requirements
  • Scope training and evaluation split rules for your use case
  • Additional formats and delivery destinations by project agreement
03Sources and delivery

Choose the coverage your model needs.

Retail datasets

Review available Home Depot and Lowe's product, price, and availability history. Store coverage and observation dates are confirmed with your sample.

Custom source lists

Amazon, Walmart, Target, marketplaces, and other public sources can be assessed as custom collection projects. Availability, permitted use, and field coverage are confirmed during scoping.

Our team prepares and delivers the data. Your client workspace is where you inspect, compare, and download it. Explore source coverage.

04Dataset context

Know what is in your training data.

Inspect available observation and delivery metadata alongside your records. Daily uploads record an observation date; they do not invent an exact capture time.

source_url
Source page associated with the observation
observed_at
Observation date or supplied capture timestamp; daily CSV imports use UTC noon
ingested_at
When the record was added to the dataset
source_run_id
Identifier for the imported batch
robots_txt_state
Recorded signal when checked; otherwise unverified
tdm_opt_out
Recorded signal when checked; null means unknown
05Questions

AI training data questions.

What formats do you deliver AI training data in?

The client portal exports CSV, JSON, and Parquet. JSONL, sharding, and other delivery requirements are scoped separately for each project.

Can you collect data from sources you do not already cover?

Send us your source list, required fields, volume, and intended use. We assess the scope and agree a sample delivery before committing to the full dataset.

How do I inspect my dataset before using it?

Use your dashboard to filter granted data by retailer, store, date, brand, and category where supplied. Inspect trends, availability, and individual products, then download the records you need.

Does source metadata establish permission to train a model?

No. Public availability and source metadata do not establish permission to train a model. We assess source restrictions and the appropriate rights basis for the intended use during scoping. Unverified or missing signals remain unknown; a project agreement cannot grant third-party rights we do not hold.

Do you provide annotation or evaluation splits?

These can be included in a custom project specification. We agree the labelling, quality checks, deduplication method, and split rules with you before preparing the delivery.

How is AI training data priced?

Pricing depends on sources, field depth, historical coverage, volume, and preparation requirements. Contact us for a sample and a project quote.

Public data.
Responsible collection.

We collect only publicly available web data. Our policy requires source review, respect for access restrictions, and a defined purpose for every project. Public access alone does not establish rights to redistribute content or use it for AI training.

Read our data collection policy

→ 03 · Begin

Tell us what data you need. We'll deliver it.

One short form. No sales decks. We'll respond within a business day with a sample, a price, and a timeline.