← Back home

2026.09 / database memory 065

Coolify PostgreSQL OOMKilled or restarting? Fix memory pressure

A PostgreSQL database in Coolify may appear healthy, disappear, recover, and then repeat the cycle under traffic, a report, an index build, or a backup. Application logs show broken connections. Coolify shows restarts. PostgreSQL may log an interrupted startup without ever writing “out of memory” itself.

Do not begin by raising every memory setting or removing the container limit. First prove which boundary failed: the PostgreSQL container’s hard limit, the whole VPS, a health-check timeout, or an ordinary database crash. Then identify the workload and setting that multiplied memory demand.

Preserve the database volume. An OOM restart is a capacity and workload incident, not a reason to delete storage, recreate PostgreSQL from scratch, or expose port 5432 publicly.

What OOMKilled actually tells you

EvidenceLikely meaningNext check
Container state says OOMKilled=trueThe container crossed a cgroup memory boundary and Docker recorded the kill.Coolify resource limit, container limit, workload at the kill time.
Kernel log names a postgres processThe host exhausted usable memory and swap, then the Linux OOM killer selected PostgreSQL.Whole-host consumers, overcommit, swap, and concurrent workloads.
Exit code or restart without OOM evidencePostgreSQL may have crashed, been stopped, failed a probe, or been recreated.Database logs, Docker events, health history, deployment activity.
Healthy container but slow or failed queriesMemory pressure may be causing swap or heavy temporary-file I/O rather than a kill.Metrics, query plans, temporary files, host I/O, and latency.

A restart count proves only that the container restarted. An unhealthy status proves only that the configured probe failed. Neither proves an OOM event. Keep the diagnosis tied to timestamps and evidence.

1. Record the incident before restarting again

In Coolify, open the PostgreSQL resource and capture its status, recent logs, metrics, current image, health state, and configured resource limits. Note the exact time of the application failure and database restart. Avoid changing several controls at once.

If you administer the host, inspect the Coolify-managed container using its actual name or ID:

docker inspect DATABASE_CONTAINER \
  --format 'oom={{.State.OOMKilled}} exit={{.State.ExitCode}} restarts={{.RestartCount}} started={{.State.StartedAt}} finished={{.State.FinishedAt}}'

docker inspect DATABASE_CONTAINER \
  --format 'memory={{.HostConfig.Memory}} reservation={{.HostConfig.MemoryReservation}} swap={{.HostConfig.MemorySwap}}'

docker logs --since 30m DATABASE_CONTAINER

Then check the host’s kernel journal for the same time window. Access to kernel logs normally requires administrative privileges:

journalctl -k --since '30 minutes ago' \
  | grep -Ei 'out of memory|oom-kill|killed process'

docker events --since '30m' \
  --filter container=DATABASE_CONTAINER

Do not publish raw logs without review. They can contain database names, role names, queries, paths, or application details. Record only the evidence needed to classify the incident.

2. Separate container OOM from host OOM

Coolify can set a maximum memory limit on the database container. Reaching that hard limit can kill PostgreSQL even while the VPS still has free memory. Conversely, a database with no explicit container limit can contribute to host-wide exhaustion and be selected by the kernel OOM killer.

Use Coolify Metrics and a short live sample to understand the normal range, but remember that a sample taken after restart misses the peak that caused the kill:

docker stats --no-stream DATABASE_CONTAINER
free -h
cat /proc/meminfo | grep -E 'MemAvailable|SwapTotal|SwapFree'

Host memory must cover more than PostgreSQL. Coolify, Docker, Traefik, application containers, builds, backups, monitoring, the kernel, and filesystem cache all share the machine. A database limit that fits on paper can still leave the VPS without safe operating headroom.

3. Read PostgreSQL’s effective memory settings

Connect through Coolify’s private database terminal or from an authorised application on the same destination network. Use the effective values, not an old configuration file or a planned environment variable:

SELECT name, setting, unit, source, pending_restart
FROM pg_settings
WHERE name IN (
  'shared_buffers',
  'work_mem',
  'hash_mem_multiplier',
  'maintenance_work_mem',
  'autovacuum_work_mem',
  'max_connections',
  'max_parallel_workers_per_gather',
  'max_parallel_maintenance_workers'
)
ORDER BY name;

These settings do not add up as one simple fixed total:

This is why setting work_mem to a large value because one report spills to disk can destabilise unrelated traffic. Several active sessions can each execute several memory-using plan nodes, and parallel workers can multiply the demand again.

4. Match the spike to database activity

Inspect session ownership, state, query start time, wait type, and parallel relationships. Avoid copying full private query text into public incident reports:

SELECT
  pid,
  leader_pid,
  datname,
  usename,
  application_name,
  state,
  wait_event_type,
  wait_event,
  backend_start,
  xact_start,
  query_start
FROM pg_stat_activity
WHERE pid <> pg_backend_pid()
ORDER BY query_start NULLS LAST, backend_start;

Look for a report, import, index build, migration, backup, autovacuum wave, worker fan-out, or new deployment that began just before the memory climb. Also compare connection counts with the deployment-wide pool budget. Too many concurrent backends can turn otherwise modest per-query settings into host pressure; the separate Coolify PostgreSQL connection-capacity guide shows how to attribute and bound those sessions.

For one currently connected backend, pg_backend_memory_contexts exposes the memory contexts of that session. An authorised administrator can also request that another backend log its memory contexts with pg_log_backend_memory_contexts(pid). Use this sparingly: it can generate many log lines and is evidence for a specific investigation, not a permanent high-frequency monitor.

5. Recover with the narrowest safe change

If the database is down, preserve the volume and configuration, stop the workload that triggered the spike, and start PostgreSQL once. Do not create a restart loop by immediately releasing the same report queue or migration.

Choose the repair that matches the evidence:

Some PostgreSQL settings reload; others require restart. Check pg_settings.context and pending_restart, make one controlled change, and let Coolify recreate or restart the resource only when required.

Do not “fix” OOM by disabling the OOM killer

Docker supports controls related to OOM behaviour, but protecting an unbounded database process can move the failure to Docker, Coolify, SSH, or another critical host process. PostgreSQL documentation also discusses Linux overcommit and OOM score adjustment, but these are host-design decisions—not first-response toggles for an undersized shared VPS.

Likewise, swap can provide a short buffer but is not substitute RAM for a busy database. Heavy swapping can turn a sharp failure into severe latency, health-check failures, and timeouts. Measure it and size the system for the workload.

6. Set Coolify limits from measured headroom

In the PostgreSQL resource, open Configuration → Resource Limits. Coolify supports a soft memory reservation, maximum memory limit, swap limit, swappiness, and CPU controls. A value of zero means no configured limit for the relevant limit fields.

A useful hard limit must satisfy two constraints:

database limit > tested peak database demand + operational burst margin

database limit + other services + builds/backups + host reserve
< safely usable host memory

Do not select a number from a generic tuning chart. Measure the actual application mix, including startup recovery, migrations, backups, autovacuum, reporting, and a representative concurrency peak. Save the Coolify limit, restart or redeploy the database as required, and confirm the running container received the intended values.

7. Verify more than a green health badge

  1. Confirm the database container starts once and its restart count remains stable.
  2. Confirm Docker no longer reports an OOM kill and the kernel records no new OOM event.
  3. Wait for PostgreSQL crash recovery, if any, to complete before judging readiness.
  4. Run a database-backed application read and one controlled write with read-back.
  5. Run the worker or report path that previously triggered the incident at bounded concurrency.
  6. Watch Coolify Metrics and host memory through the full operation, not only its first seconds.
  7. Confirm connection counts remain within the planned pool budget.
  8. Run a fresh scheduled backup and verify completion.
  9. Check that application latency, temporary-file I/O, and swap activity remain acceptable.

A passing PostgreSQL health check proves that its configured probe succeeded. It does not prove that a representative query fits the memory budget, a backup can complete, or the application can read and write safely.

Compact incident order

  1. Record the failure time, container state, restart count, logs, metrics, and Coolify limits.
  2. Prove container OOM, host OOM, or a different exit cause.
  3. Preserve the persistent volume and stop the triggering workload before recovery.
  4. Read effective PostgreSQL memory, connection, maintenance, and parallel-worker settings.
  5. Match the spike to queries, jobs, migrations, maintenance, backups, or a deployment.
  6. Reduce the narrow multiplier: query cost, concurrency, pool size, parallelism, or an oversized setting.
  7. Raise the Coolify limit or host capacity only from measured demand and real headroom.
  8. Restart once when required, then verify reads, writes, jobs, backups, memory, and restart stability.

The durable fix is not “give PostgreSQL all the RAM.” It is to make the memory boundary and workload multipliers explicit. Coolify controls the container boundary; PostgreSQL controls shared and per-operation memory; the application controls concurrency; and every service still shares one host. When those budgets agree, an analytical query or routine backup is far less likely to become a database outage.

Authoritative references: Coolify documents database resource limits, database health checks, and its standalone database model. Docker explains container memory constraints and OOM risk. PostgreSQL documents shared, query, maintenance, and parallel-worker memory, Linux memory overcommit and OOM behaviour, pg_stat_activity, and backend memory-context logging.