All case studies
Case study

Making cache synchronisation resilient with SQS, retries and a DLQ

Decoupling PostgreSQL writes from Redis updates with an SQS worker, retries, a DLQ and CloudWatch alarms, so cache failures stop affecting the API.

Context
Backend of a high-traffic consumer platform (client under NDA)
Role
Tech Lead / Backend
Key point
Cache sync off the request path
  • PostgreSQL
  • Redis
  • AWS SQS
  • DLQ
  • CloudWatch
  • SNS
  • Go / Node.js workers

The problem

Each write updated PostgreSQL and then Redis inside the same request.

The API had to wait for the cache update, a Redis failure could fail the user's request, and retry logic was scattered through the code.

Architecture: before → after

Before
  1. Database write
  2. Inline Redis update
  3. Tightly coupled, failures leak to the API
After
  1. Database write
  2. Publish SQS event
  3. Worker updates Redis
  4. Retry → DLQ → CloudWatch → SNS

What I did

  • Kept the database write in the request, and published an event to SQS instead of updating Redis inline.
  • A worker consumes the queue and updates Redis.
  • Failed messages are retried and, after repeated failures, routed to a dead-letter queue.
  • CloudWatch metrics and alarms on the queues, with SNS notifications, make failed synchronisation visible.

Outcome

  • The API path became less dependent on Redis availability.
  • Cache synchronisation became asynchronous, with retries handled in one place.
  • Failed synchronisation became observable instead of silent.