Skip to main content
All articles

The HPA Connection Tax: Why Horizontal Scaling Is Crashing Your Database

Autoscaled pods can exhaust database connections and cascade into an outage. Pooling and scaling guardrails that prevent it.

Sahil BansalConnect

4 min readOriginally on Medium

When traffic spikes, the standard response is to let the Horizontal Pod Autoscaler (HPA) do its job. But in many production environments, HPA acts as a force multiplier for database failure, accelerating connection exhaustion and triggering cascading outages across the stateful tier.

Why This Matters

We build distributed systems to handle load, yet we often ignore the math of stateful bottlenecks. In a typical microservices architecture, scaling from 10 to 100 pods might solve a CPU bottleneck but immediately hits the wall of Postgres or MySQL connection limits.

Most SREs treat the database as a fixed resource, but HPA treats compute as infinitely elastic. When these two philosophies collide during a traffic surge, you don’t just get latency; you get a complete system blackout as the database begins rejecting all incoming traffic from both old and new pods.

Architecture / System Design

In a standard Kubernetes deployment, each pod maintains its own internal connection pool. This is designed for efficiency at the application level but is disastrous for global resource management.

When HPA triggers a scale-up event, it launches new replicas. Each replica immediately attempts to reserve its minimum_idle or initial_size connections. In a high-churn environment, the database spends more CPU cycles managing connection handshakes and process forking than executing queries.

If your database is configured with max_connections = 2000 and your HPA is configured to scale up to 100 replicas with a max_pool_size of 30, you have a hard ceiling at 66 replicas. Replicas 67 through 100 will fail to start, but worse, their constant retry loops will saturate the database's listener thread, slowing down existing healthy connections.

Implementation

To prevent HPA from becoming a self-inflicted DDoS, you must decouple pod count from connection count. This is best achieved through a connection multiplexer like PgBouncer for Postgres or RDS Proxy for AWS environments.

PgBouncer Sidecar Configuration

Instead of connecting directly to the DB, the application connects to a local sidecar. This allows the application to maintain a “warm” pool locally while the sidecar manages a much smaller, shared pool to the actual database.

apiVersion: apps/v1
kind: Deployment
metadata:
 name: order-service
spec:
 template:
    spec:
      containers:
      - name: app
        env:
        - name: DB_HOST
          value: "localhost"
        - name: DB_PORT
          value: "6432"
      - name: pgbouncer
        image: edoburu/pgbouncer
        ports:
        - containerPort: 6432
        env:
        - name: DATABASE_URL
          value: "postgres://user:pass@prod-db.cluster:5432/orders"
        - name: PGBOUNCER_POOL_MODE
          value: "transaction"
        - name: PGBOUNCER_MAX_CLIENT_CONN
          value: "100"
        - name: PGBOUNCER_DEFAULT_POOL_SIZE
          value: "10"

HPA Behavior Guardrails

Your HPA must be configured with a scaleDown stabilization window to prevent "flapping," which causes constant connection churn.

apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
 name: order-service-hpa
spec:
 behavior:
    scaleDown:
      stabilizationWindowSeconds: 300
      policies:
      - type: Percent
        value: 10
        periodSeconds: 60
    scaleUp:
      stabilizationWindowSeconds: 0
      policies:
      - type: Pods
        value: 5
        periodSeconds: 30

Operational Realities

Scaling behavior is rarely linear. As you add pods, the overhead of the database managing those connections grows. In Postgres, each connection is a separate process. At high connection counts, context switching between these processes becomes a significant percentage of total DB CPU usage.

Observability must focus on the connection_utilization metric. If you see this spike concurrently with HPA scaling events, you are under-provisioned at the proxy layer. You need to monitor the time spent in connection_acquisition within the application; if this rises, your sidecar is saturated or the DB is bottlenecked on process forking.

Deployment risk is also higher. A bad deployment that causes pods to crash-loop can exhaust connections instantly if pods don’t close connections gracefully on SIGTERM. Always implement a preStop hook to ensure the application pool shuts down before the container is killed.

Failure Modes / Trade-offs

Using a transaction-level proxy (like PgBouncer in transaction mode) introduces specific limitations. You cannot use session-based features like SET TIME ZONE or prepared statements if the proxy doesn't support them. This forces developers to write more "pure" SQL, which is a trade-off for scalability.

Another failure mode is the “Thundering Herd” at the proxy itself. If the proxy’s own incoming connection limit is too low, the HPA replicas will bottleneck there instead of the DB.

Cost is another factor. RDS Proxy and similar managed services are not free. However, the cost of an RDS Proxy is almost always lower than the cost of over-provisioning a massive DB instance just to handle the memory overhead of thousands of idle connections.

Lessons Learned / Best Practices

  • Math First: Always calculate HPA_Max * App_Pool_Size and ensure it is < DB_Max_Connections * 0.8.
  • Transaction Pooling: Use transaction-level pooling instead of session-level pooling to maximize connection reuse.
  • Graceful Shutdown: Ensure your application handles SIGTERM by draining its connection pool, preventing "zombie" connections on the DB.
  • Proxy Everything: Never let an autoscaling service talk directly to a database. Use a proxy layer to buffer the connection surges.

TL;DR

  • HPA is infrastructure-blind; it scales based on local CPU/Memory metrics without awareness of downstream database connection limits.
  • Application-level connection pooling (e.g., HikariCP, SQLAlchemy) becomes a liability when pod counts fluctuate rapidly.
  • Total connections scale linearly with pods: Total = (HPA_Max_Replicas * Max_Pool_Size). If this exceeds the DB’s max_connections, the last 20% of your pods will kill the first 80%.
  • Sidecar-based or centralized proxying (PgBouncer, RDS Proxy) is mandatory for HPA-driven workloads, not optional.
  • Connection churn during rapid scale-up adds significant latency to startup times, often triggering liveness probe failures.

Final Thoughts

Kubernetes makes it easy to scale compute, but it doesn’t solve the fundamental constraints of stateful resources. If your HPA configuration doesn’t account for the connection limits of your database, you haven’t built a self-healing system — you’ve built a system that is engineered to fail faster under pressure.

  • Kubernetes
  • Autoscaling
  • PostgreSQL

Written by Sahil Bansal

DevOps and platform engineer. I write about the infrastructure decisions I have had to live with.

Connect