16 August 2016, Amazon. Project

Owned the team's move of 57 hosts to Amazon Linux

I owned the team's migration off an end-of-life operating system across 57 hosts, 22 environments and 9 pipelines, kept the feeds org informed on one tracking page, and closed the policy violation and its open tickets.

  • Reliability
  • Delivery

57

hosts replaced

22

environments

alpha, beta, gamma and prod

9

pipelines

What would have happened

Hosts stayed on an unsupported OS, the policy-engine violation stayed open and the ticket kept showing up in weekly SLA reports.

The call

Put availability first. After one environment failed, move one host at a time, beta before prod, and stop on any failure.

What I did

Tracked every environment on one page, cleared two blockers, and restored beta and prod capacity when one environment failed mid-move before carrying on.

What changed

57 hosts, 22 environments and 9 pipelines moved, 4 related tickets closed, and my manager replied 'Fantastic. Nice work.'

Proof

The original documents are held in my private record and can be shared on request.

What people said

“Fantastic. Nice work. Thanks for the update.”

Engineering managerAmazon, 18 August 2016