Built the data pipeline behind the catalog value reports, about 2 billion records
To measure how much of the catalog came from automated feed parsing, I built a tool that pulled catalog search results, split them and loaded them into the data warehouse, back to 1970 across 12 marketplaces.
- Catalog
- Delivery
~2B
catalog records backfilled
930M vendor-contribution plus 1.1B create-by-date responses, back to 1970
12
marketplaces
~30,000
items processed per day
ongoing daily job
What would have happened
Value-metrics reports on automated feed parsing would have run on incomplete history, because the warehouse could not ingest hundreds of gigabytes of single-line responses.
The call
Automate the pull, split the data in code and backfill all history, then keep it current with daily jobs.
What I did
Designed and wrote the query generator, the splitter and uploader, and set up the daily jobs that keep the data current.
What changed
About 2 billion records were backfilled and about 30,000 new items flow in each day, so the reports cover the full history.
Proof
The original documents are held in my private record and can be shared on request.
