Image Resize Service Performance Degradation

Incident Report for RebelMouse

Postmortem

Summary

Between 14:14 and 14:30 UTC, the Image Resize service was unable to keep up with incoming demand, causing a portion of asset requests to fail with HTTP 503 responses. The service was saturated by a sharp, sustained increase in crawler traffic that pushed request volume well beyond the capacity provisioned for normal baseline load.

The incident was mitigated by significantly increasing the resources and capacity available to the service. Error rates began falling within minutes of the capacity change and returned to baseline by 14:30 UTC. No data loss occurred, and no customer content was affected — failures were limited to transient delivery errors on individual asset requests.

Impact

  • Failed requests: 38,248 asset 503 responses during 14:14-14:30 UTC timeframe.

  • Total requests served: approximately 1.31 million asset requests in the same interval.

  • Aggregate error rate: approximately 2.91%.

Roughly 97% of asset requests continued to succeed throughout the incident, and users retrying a failed image load would frequently have received it successfully.

Timeline

  • 14:14 Elevated 503 error rates began on asset requests.
  • 14:20 Peak impact: 5.65% of asset requests failing.
  • 14:25 Additional capacity took effect; error rates began declining.
  • 14:30 Incident is resolved.

Root cause

The Image Resize service experienced a large increase in load driven by crawler traffic. Crawlers requesting image assets generated a volume of resize work substantially higher than typical baseline demand, and heavily weighted toward uncached variants, which meant a large share of that traffic had to be served by origin resize workers rather than from cache.

Mitigation and resolution

Once the resize tier was identified as the saturation point, available resources and capacity for the service were increased significantly. Error rates began declining at 14:25, fell steadily over the following four minutes, and returned to baseline by 14:30. Normal behavior was confirmed from 14:30 UTC onward.

The expanded capacity was left in place after recovery pending a review of steady-state sizing.

Action items

Review and permanently raise baseline capacity for the resize tier to restore adequate burst headroom.

Posted Aug 26, 2026 - 08:06 EDT

Resolved

We are currently investigating this issue.
Posted Aug 25, 2026 - 10:30 EDT