Why this lesson exists
Most new PostgreSQL in companies runs as a managed service: Azure Database for PostgreSQL (Flexible Server), Amazon RDS for PostgreSQL and Aurora PostgreSQL, Google Cloud SQL and AlloyDB. The provider installs, patches, backs up and fails over - and teams conclude that "the database is the cloud's problem". It is not. Every incident in this chapter still happens on a managed server: too many clients, a blocking chain, autovacuum held back by an idle transaction, a slow query, a migration that takes the site down, a replication slot filling the disk. What changes is which knobs you have and where you look. This lesson maps the chapter onto the managed world.
What you need to know already: this whole chapter, and the Azure chapters (resource groups, VNets, private endpoints, managed identities, Key Vault).
The words you need first
- Shared responsibility - the provider owns the hardware, the OS, the PostgreSQL binaries, minor patching and the backup machinery; you own the schema, queries, connections, roles, settings within what is allowed, capacity, and testing that restores and failovers actually work for you.
- Server parameters - the managed replacement for
postgresql.conf: a portal page / CLI / Terraform resource. Static parameters still need a restart, which the service performs. - No superuser - you get an admin role with most privileges (
azure_pg_admin,rds_superuser,cloudsqlsuperuser) but notSUPERUSER: no server filesystem, noCOPY ... TO '/file', no arbitrary extensions, noALTER SYSTEM. - HA - a standby in another zone with automatic failover, behind one DNS name.
- PITR window - the retention of automated backups + WAL (7 to 35 days); restores create a new server.
What is the same, what changed
| chapter topic | on a managed server |
|---|---|
| install, files, service (lesson 2) | gone: no SSH, no pg_lsclusters, no systemctl. Logs go to the provider's log service (Azure Monitor / Log Analytics, CloudWatch Logs, Cloud Logging) |
| config, reload vs restart | server parameters via portal/CLI/IaC; "dynamic" (reload) vs "static" (restart) - the service tells you which; some are locked |
| psql (lesson 3) | the same - over TLS from a jump host, a bastion or a private network; sslmode=verify-full with the provider's CA |
| roles, pg_hba (lesson 4) | no pg_hba.conf: network rules (firewall, VNet, private endpoint, security groups) + roles. Passwords or cloud identities (Microsoft Entra ID tokens, IAM database authentication) - no long-lived passwords at all |
| connections, PgBouncer (lesson 5) | same arithmetic, max_connections often derived from the instance size. Azure Flexible Server has a built-in PgBouncer (port 6432, server parameter pgbouncer.enabled); AWS has RDS Proxy; Cloud SQL has a managed connection pooler |
| locks, MVCC (lesson 6) | identical: pg_stat_activity, pg_blocking_pids, pg_terminate_backend all work (the admin role may signal non-superuser sessions) |
| vacuum, wraparound (lesson 7) | identical - and just as easy to break with a long transaction. The providers alert on XID age and may force maintenance |
| EXPLAIN, pg_stat_statements (8, 9) | the same SQL; pg_stat_statements usually preloaded. Extra tools: Query Store / Query Performance Insight (Azure), Performance Insights / Database Insights (AWS), Query Insights (GCP) |
| migrations (lesson 10) | identical risk, identical fixes (lock_timeout, CONCURRENTLY) |
| backups (11, 12) | automated daily snapshots + continuous WAL = PITR to any second in the window, as a new server. Logical dumps (pg_dump, \copy) still yours, and still worth having |
| replication, failover (lesson 13) | HA standby (zone-redundant), failover in about a minute behind the same DNS name; read replicas (async, possibly cross-region) you create and promote. Slots for logical replication / CDC can still fill storage |
| monitoring (lesson 14) | the provider's metrics (CPU, storage, connections, replication lag) + your own: postgres_exporter still works against a managed server, with a pg_monitor role |
What you still own
- Connections. The managed server enforces
max_connectionsthe same way, and the reserved slots belong to the provider's own agents. Pool size x replicas, PgBouncer in transaction mode, per-role limits. - Credentials. Prefer identity-based auth: an app with a managed identity (Azure) or an IAM role (AWS) gets a short-lived token instead of a password. If passwords stay, they live in Key Vault / Secrets Manager and rotate - or Vault issues them dynamically (the Vault chapter).
- Network exposure. Private access (VNet integration / private endpoint / private IP) and no public endpoint for production; TLS enforced (
require_secure_transport/rds.force_ssl). - Capacity. Storage auto-grow exists but costs money and only grows; IOPS and throughput are tied to the tier or provisioned separately - a "slow database" is often an I/O limit on the disk tier. Watch storage used, IOPS and throughput limits, CPU credits on burstable tiers.
- Maintenance windows. Minor version patches and host maintenance restart the server at a time you pick - make sure the apps reconnect (and fail over) cleanly. Major upgrades are still a project: test, in-place upgrade or a blue/green switch (RDS Blue/Green, logical replication).
- Restore and failover drills. "Restore to a point in time" makes a new server with a new name; your runbook needs to say how the apps get pointed at it, and how long it takes for your data size (RTO). Trigger a planned failover once a quarter and watch what the apps do.
- The data. Schema, indexes, queries, vacuum health, bloat, retention - exactly as on your own VM.
Choosing between them
| Azure Database for PostgreSQL Flexible Server | Amazon RDS for PostgreSQL | Amazon Aurora PostgreSQL | |
|---|---|---|---|
| what it is | community PostgreSQL on Azure VMs + managed storage | community PostgreSQL on EC2 + EBS | PostgreSQL-compatible engine on a distributed storage layer |
| HA | zone-redundant standby (sync), ~60-120 s failover | Multi-AZ standby (sync); Multi-AZ cluster with 2 readable standbys | up to 15 replicas on shared storage, failover typically under a minute |
| pooling | built-in PgBouncer | RDS Proxy (separate, billed) | RDS Proxy |
| auth | passwords and/or Microsoft Entra ID | passwords and/or IAM auth | passwords and/or IAM auth |
| backups | automated, PITR 7-35 days, geo-redundant option | automated snapshots, PITR up to 35 days | continuous, PITR, backtrack (rewind in place, MySQL only) |
| versions | majors usually within months of release | same | trails community releases more |
Rules of thumb: on Azure, Flexible Server is the default (Single Server was retired in 2025). On AWS, RDS when you want plain PostgreSQL and predictable costs; Aurora when you need many read replicas, very fast failover or storage that grows to large sizes without planning. On Kubernetes, an operator (CloudNativePG) gives you most of this inside the cluster - and makes you the provider again.
The checks you ran by hand in this chapter work against all of them. A quick health query you can run on any PostgreSQL, managed or not:
$ sudo -u postgres psql -c "select (select count(*) from pg_stat_activity where backend_type = 'client backend') as conns, current_setting('max_connections') as max_conns, (select max(now() - xact_start) from pg_stat_activity) as oldest_xact, (select count(*) from pg_stat_activity where state = 'idle in transaction') as idle_in_tx, (select max(age(datfrozenxid)) from pg_database) as max_xid_age, (select count(*) from pg_replication_slots where not active) as inactive_slots"
conns | max_conns | oldest_xact | idle_in_tx | max_xid_age | inactive_slots
-------+-----------+-------------+------------+-------------+----------------
1 | 100 | | 0 | 80 | 0
(1 row)
In an interview: "We moved to managed PostgreSQL - what is still our job?" - connections and pooling, credentials (prefer managed identities / IAM tokens, rotation), network exposure (private endpoints, TLS), capacity (storage, IOPS limits), the schema/queries/vacuum health, safe migrations, and testing restores and failovers - including how apps get pointed at a restored server and how they reconnect after a failover. The provider runs the machinery; you own whether the service survives it.
What you can do now
- Map each part of this chapter onto a managed service: what disappears, what changes, what stays.
- Name what a managed service does not do for you, and plan for it.
- Compare Azure Flexible Server, RDS and Aurora on HA, pooling, auth and backups.
- Run the same health checks on any PostgreSQL, managed or self-hosted.