January 2018, Amazon. Project

Cut garbage collection to 2.4% of CPU on a fraud service

I profiled a fraud-prevention service, found it spent about 17.7% of CPU on garbage collection, tuned the JVM to bring that to 2.4%, and showed the team the before-and-after evidence.

  • Reliability
  • Cost
  • Fraud and risk

~17.7% to ~2.4%

CPU spent on garbage collection

~20%

more supported TPS

What would have happened

More hosts to handle peak load, with CPU burned on non-business work.

The call

Measure with the profiler before changing code or buying capacity.

What I did

Read the flame graphs, raised the heap size and moved to the G1 collector, then posted before-and-after charts to the team with notes on how to read them.

What changed

Garbage collection dropped from about 17.7% to about 2.4% of CPU, and supported throughput rose about 20%.

Proof

The original documents are held in my private record and can be shared on request.