# System Design: How to Handle High Traffic Without Breaking Your App
Why High Traffic Scares Backend Engineers
Let’s be honest—when traffic spikes, backend engineers lose sleep. Not because we can’t handle it, but because we’ve seen what happens when things go wrong. A sudden surge from a viral post, a bot attack, or a misconfigured cron job can turn a stable system into a dumpster fire in minutes.
I’ve been there. In 2023, one of our services at a fintech startup got hit with 10x traffic after a feature went viral on Twitter. Our Postgres primary couldn’t keep up, Redis evictions started kicking in, and our Node.js servers were throwing ECONNRESET errors like confetti. We survived, but it wasn’t pretty.
The goal isn’t to build a system that *never* fails—it’s to build one that fails *gracefully* and recovers fast. Here’s how real teams do it.
---
Load Balancing: Don’t Let One Server Take the Hit
If you’re running a single server, you’re doing it wrong. Even a small app needs at least two instances behind a load balancer. Why? Because hardware fails, deployments break, and one instance can’t handle 10k RPS alone.
How It Works in Production
- Round Robin: Simple, works for stateless apps. But if one instance is slow, users get stuck with latency.
- Least Connections: Better for long-lived connections (WebSockets, gRPC). Avoids overloading a single instance.
- IP Hash: Ensures a user always hits the same server (useful for session stickiness).
Practical Example: Nginx Config
upstream backend {
least_conn;
server 10.0.1.1:3000;
server 10.0.1.2:3000;
server 10.0.1.3:3000;
}
server {
listen 80;
location / {
proxy_pass http://backend;
proxy_set_header Host $host;
}
}
Tradeoffs
- Pros: Scales horizontally, improves uptime.
- Cons: Adds latency (~1-5ms), needs health checks, session management becomes tricky.
Failure Case: We once had a load balancer misconfigured to send traffic to a decommissioned instance. Took us 20 minutes to notice because the health checks were too lenient. Lesson: Tighten health check timeouts and monitor 5xx errors.
---
Horizontal Scaling: More Machines, Less Pain
Vertical scaling (bigger servers) works until it doesn’t. At some point, you hit hardware limits, and downtime becomes inevitable for upgrades.
Horizontal scaling is the answer—add more machines, distribute the load. But it’s not free.
What Scales Horizontally?
✅ Stateless Services: APIs, microservices, frontend servers. ✅ Databases with Read Replicas: Postgres, MySQL (but watch replication lag). ✅ Caches: Redis, Memcached. ✅ Queues: RabbitMQ, Kafka, SQS.
What Doesn’t Scale Horizontally?
❌ Monolithic Databases: Your Postgres primary can’t scale infinitely. ❌ Stateful Services: WebSocket connections, long-running transactions. ❌ Single Points of Failure: Cron jobs, legacy services.
Example: Auto-Scaling in AWS
# AWS CLI command to scale based on CPU
aws application-autoscaling register-scalable-target \
--service-namespace ecs \
--resource-id service/my-cluster/my-service \
--scalable-dimension ecs:service:DesiredCount \
--min-capacity 2 \
--max-capacity 20
Failure Case: We auto-scaled too aggressively once, and a bug in our app caused a feedback loop—more traffic → more instances → more traffic. Ended up with 50 instances and a $12k AWS bill. Lesson: Set conservative max limits and monitor scaling events.
---
Caching: The Lazy Developer’s Best Friend
Caching is the cheapest way to reduce database load. But it’s also the easiest way to introduce bugs.
Where to Cache
- CDN: Static assets (images, JS, CSS).
- Redis/Memcached: API responses, computed data (e.g., leaderboards).
- Database Query Cache: Postgres
pg_cache, MySQL query cache (but be careful—it’s often disabled for a reason).
Example: Node.js + Redis Cache
const redis = require('redis');
const client = redis.createClient({ url: 'redis://redis-cache:6379' });
async function getUserProfile(userId) {
const cacheKey = `user:${userId}`;
const cached = await client.get(cacheKey);
if (cached) {
return JSON.parse(cached);
}
const user = await db.query('SELECT * FROM users WHERE id = $1', [userId]);
await client.set(cacheKey, JSON.stringify(user), 'EX', 300); // Cache for 5 mins
return user;
}
Cache Invalidation: The Hard Part
- Time-based (TTL): Simple, but stale data can cause issues.
- Event-based: Invalidate cache on writes (e.g.,
user:123→ update → delete cache). - Write-through: Update cache on every write (can cause thundering herd).
Failure Case: We cached user profiles with a 1-hour TTL. A bug in our admin panel updated a user’s email but didn’t invalidate the cache. Users saw old emails for an hour. Lesson: Always invalidate cache on writes, even if it’s slower.
---
Read Replicas: Let Others Do the Heavy Lifting
If your database is the bottleneck, read replicas are the first line of defense. But they come with tradeoffs.
When to Use Read Replicas
- Read-heavy workloads: Analytics, dashboards, product listings.
- High-latency queries: Reports, aggregations.
- Disaster recovery: Failover to a replica if primary dies.
When NOT to Use Read Replicas
- Write-heavy workloads: Replicas won’t help if your primary is struggling.
- Strong consistency needs: Replicas can lag (seconds to minutes).
- Frequent schema changes: Replication breaks can happen.
Example: Postgres Read Replica
-- Connect to replica (read-only)
SHOW transaction_read_only; -- Should return "on"
-- Application code should route reads to replicas
const isReadQuery = req.method === 'GET';
const db = isReadQuery ? replicaPool : primaryPool;
Failure Case: We once had a replica fall behind by 30 minutes during a traffic spike. Our dashboard showed stale data, and users complained. Lesson: Monitor replication lag (pg_stat_replication) and set alerts.
---
Queues: Don’t Do Heavy Work in Requests
If your API does anything slow (emails, payments, PDF generation), offload it to a queue.
Why Queues Help
- Decoupling: API doesn’t wait for slow tasks.
- Retry Logic: Failed jobs can retry.
- Rate Limiting: Control how fast tasks execute.
Example: Node.js + Bull Queue
const Queue = require('bull');
const emailQueue = new Queue('email', 'redis://redis-queue:6379');
app.post('/send-email', async (req, res) => {
await emailQueue.add({
to: req.body.email,
subject: 'Welcome!',
body: 'Thanks for signing up.'
});
res.send('Email queued!');
});
// Worker process
emailQueue.process(async (job) => {
await sendEmail(job.data);
});
Queue Pitfalls
- Backpressure: If jobs pile up, your queue becomes a bottleneck.
- At-least-once delivery: Jobs might run twice (design for idempotency).
- Dead letters: Failed jobs need a place to go.
Failure Case: We once had a queue worker crash silently, and jobs piled up for hours. By the time we noticed, the Redis queue was 10GB. Lesson: Monitor queue length and worker health.
---
Rate Limiting: Stop Abuse Before It Starts
Rate limiting isn’t just for security—it’s for survival. Without it, a single script kiddie can take down your API.
Where to Rate Limit
- API Endpoints:
/login,/payments,/search. - Database Queries: Prevent expensive queries from running too often.
- Third-Party APIs: Don’t get rate-limited by Stripe or Twilio.
Example: Redis + Rate Limiting
const { RateLimiterRedis } = require('rate-limiter-flexible');
const redisClient = require('redis').createClient();
const rateLimiter = new RateLimiterRedis({
storeClient: redisClient,
keyPrefix: 'rate_limit',
points: 10, // 10 requests
duration: 1, // per 1 second
});
app.use(async (req, res, next) => {
try {
await rateLimiter.consume(req.ip);
next();
} catch (err) {
res.status(429).send('Too Many Requests');
}
});
Rate Limiting Strategies
- Fixed Window: Simple, but can allow bursts at window edges.
- Sliding Window: More accurate, but harder to implement.
- Token Bucket: Good for APIs with bursty traffic.
Failure Case: We didn’t rate-limit /search initially. A bot started scraping our site at 100 RPS, causing database CPU to spike to 90%. Lesson: Rate limit everything by default.
---
CDN: Offload the Easy Stuff
CDNs (Cloudflare, Fastly, Akamai) are the easiest way to reduce server load. They cache static assets and even dynamic content.
What to Cache with a CDN
- Static Files: JS, CSS, images.
- API Responses: If they’re cacheable (e.g.,
/products). - HTML Pages: For marketing sites, blogs.
Example: Cloudflare Cache Rules
Cache Level: Cache Everything
Edge Cache TTL: 1 hour
Browser Cache TTL: 1 day
CDN Pitfalls
- Cache Invalidation: Hard to do at scale.
- Dynamic Content: CDNs can’t cache everything.
- Cost: Bandwidth isn’t free.
Failure Case: We cached /api/products for 1 hour, but a pricing update didn’t reflect for users. Lesson: Use shorter TTLs for dynamic data or purge cache on updates.
---
Database Pressure: The Silent Killer
Databases are the hardest thing to scale. If you don’t optimize queries, indexes, and connections, you’re dead.
How to Reduce Database Load
- Index Smartly: But don’t over-index—writes get slower.
- Connection Pooling: Don’t open/close connections per request.
- Read Replicas: Offload reads.
- Batch Writes: Combine inserts/updates.
- Archiving: Move old data to cold storage.
Example: Node.js Connection Pooling
const { Pool } = require('pg');
const pool = new Pool({
max: 20, // Max connections
idleTimeoutMillis: 30000, // Close idle connections
});
app.get('/users', async (req, res) => {
const { rows } = await pool.query('SELECT * FROM users LIMIT 100');
res.json(rows);
});
Database Failure Modes
- Connection Exhaustion: Too many connections →
too many clients. - **Lock