16 June 2016, Amazon. Project

Built the data pipeline behind the catalog value reports, about 2 billion records

To measure how much of the catalog came from automated feed parsing, I built a tool that pulled catalog search results, split them and loaded them into the data warehouse, back to 1970 across 12 marketplaces.

  • Catalog
  • Delivery

~2B

catalog records backfilled

930M vendor-contribution plus 1.1B create-by-date responses, back to 1970

12

marketplaces

~30,000

items processed per day

ongoing daily job

What would have happened

Value-metrics reports on automated feed parsing would have run on incomplete history, because the warehouse could not ingest hundreds of gigabytes of single-line responses.

The call

Automate the pull, split the data in code and backfill all history, then keep it current with daily jobs.

What I did

Designed and wrote the query generator, the splitter and uploader, and set up the daily jobs that keep the data current.

What changed

About 2 billion records were backfilled and about 30,000 new items flow in each day, so the reports cover the full history.

Proof

The original documents are held in my private record and can be shared on request.