2026.09 / database memory 065
Coolify PostgreSQL OOMKilled or restarting? Fix memory pressure
A PostgreSQL database in Coolify may appear healthy, disappear, recover, and then repeat the cycle under traffic, a report, an index build, or a backup. Application logs show broken connections. Coolify shows restarts. PostgreSQL may log an interrupted startup without ever writing “out of memory” itself.
Do not begin by raising every memory setting or removing the container limit. First prove which boundary failed: the PostgreSQL container’s hard limit, the whole VPS, a health-check timeout, or an ordinary database crash. Then identify the workload and setting that multiplied memory demand.
Preserve the database volume. An OOM restart is a capacity and workload incident, not a reason to delete storage, recreate PostgreSQL from scratch, or expose port 5432 publicly.
What OOMKilled actually tells you
| Evidence | Likely meaning | Next check |
|---|---|---|
Container state says OOMKilled=true | The container crossed a cgroup memory boundary and Docker recorded the kill. | Coolify resource limit, container limit, workload at the kill time. |
Kernel log names a postgres process | The host exhausted usable memory and swap, then the Linux OOM killer selected PostgreSQL. | Whole-host consumers, overcommit, swap, and concurrent workloads. |
| Exit code or restart without OOM evidence | PostgreSQL may have crashed, been stopped, failed a probe, or been recreated. | Database logs, Docker events, health history, deployment activity. |
| Healthy container but slow or failed queries | Memory pressure may be causing swap or heavy temporary-file I/O rather than a kill. | Metrics, query plans, temporary files, host I/O, and latency. |
A restart count proves only that the container restarted. An unhealthy status proves only that the configured probe failed. Neither proves an OOM event. Keep the diagnosis tied to timestamps and evidence.
1. Record the incident before restarting again
In Coolify, open the PostgreSQL resource and capture its status, recent logs, metrics, current image, health state, and configured resource limits. Note the exact time of the application failure and database restart. Avoid changing several controls at once.
If you administer the host, inspect the Coolify-managed container using its actual name or ID:
docker inspect DATABASE_CONTAINER \
--format 'oom={{.State.OOMKilled}} exit={{.State.ExitCode}} restarts={{.RestartCount}} started={{.State.StartedAt}} finished={{.State.FinishedAt}}'
docker inspect DATABASE_CONTAINER \
--format 'memory={{.HostConfig.Memory}} reservation={{.HostConfig.MemoryReservation}} swap={{.HostConfig.MemorySwap}}'
docker logs --since 30m DATABASE_CONTAINER
Then check the host’s kernel journal for the same time window. Access to kernel logs normally requires administrative privileges:
journalctl -k --since '30 minutes ago' \
| grep -Ei 'out of memory|oom-kill|killed process'
docker events --since '30m' \
--filter container=DATABASE_CONTAINER
Do not publish raw logs without review. They can contain database names, role names, queries, paths, or application details. Record only the evidence needed to classify the incident.
2. Separate container OOM from host OOM
Coolify can set a maximum memory limit on the database container. Reaching that hard limit can kill PostgreSQL even while the VPS still has free memory. Conversely, a database with no explicit container limit can contribute to host-wide exhaustion and be selected by the kernel OOM killer.
- Container-bound incident: Docker reports
OOMKilled=true, the database limit is near the observed peak, and the host retained headroom. - Host-bound incident: kernel logs show global memory exhaustion, several services were under pressure, and PostgreSQL may or may not have had a Coolify limit.
- Not yet proven: the database restarted but neither Docker state nor kernel logs confirm OOM. Investigate the PostgreSQL exit and health check instead of assuming.
Use Coolify Metrics and a short live sample to understand the normal range, but remember that a sample taken after restart misses the peak that caused the kill:
docker stats --no-stream DATABASE_CONTAINER
free -h
cat /proc/meminfo | grep -E 'MemAvailable|SwapTotal|SwapFree'
Host memory must cover more than PostgreSQL. Coolify, Docker, Traefik, application containers, builds, backups, monitoring, the kernel, and filesystem cache all share the machine. A database limit that fits on paper can still leave the VPS without safe operating headroom.
3. Read PostgreSQL’s effective memory settings
Connect through Coolify’s private database terminal or from an authorised application on the same destination network. Use the effective values, not an old configuration file or a planned environment variable:
SELECT name, setting, unit, source, pending_restart
FROM pg_settings
WHERE name IN (
'shared_buffers',
'work_mem',
'hash_mem_multiplier',
'maintenance_work_mem',
'autovacuum_work_mem',
'max_connections',
'max_parallel_workers_per_gather',
'max_parallel_maintenance_workers'
)
ORDER BY name;
These settings do not add up as one simple fixed total:
shared_buffersis a server-wide shared allocation established at startup.work_memis a base limit for each sort or hash operation, not one allowance for the whole server or even one allowance per query.hash_mem_multipliercan let hash operations exceed the basework_memamount.- Parallel query workers are separate processes, and relevant resource limits can apply to each worker as well as the leader.
maintenance_work_memis used by maintenance work such as index creation and some vacuum operations.- Autovacuum workers, connections, extensions, caches, query state, and the operating system consume additional memory.
This is why setting work_mem to a large value because one report spills to disk can destabilise unrelated traffic. Several active sessions can each execute several memory-using plan nodes, and parallel workers can multiply the demand again.
4. Match the spike to database activity
Inspect session ownership, state, query start time, wait type, and parallel relationships. Avoid copying full private query text into public incident reports:
SELECT
pid,
leader_pid,
datname,
usename,
application_name,
state,
wait_event_type,
wait_event,
backend_start,
xact_start,
query_start
FROM pg_stat_activity
WHERE pid <> pg_backend_pid()
ORDER BY query_start NULLS LAST, backend_start;
Look for a report, import, index build, migration, backup, autovacuum wave, worker fan-out, or new deployment that began just before the memory climb. Also compare connection counts with the deployment-wide pool budget. Too many concurrent backends can turn otherwise modest per-query settings into host pressure; the separate Coolify PostgreSQL connection-capacity guide shows how to attribute and bound those sessions.
For one currently connected backend, pg_backend_memory_contexts exposes the memory contexts of that session. An authorised administrator can also request that another backend log its memory contexts with pg_log_backend_memory_contexts(pid). Use this sparingly: it can generate many log lines and is evidence for a specific investigation, not a permanent high-frequency monitor.
5. Recover with the narrowest safe change
If the database is down, preserve the volume and configuration, stop the workload that triggered the spike, and start PostgreSQL once. Do not create a restart loop by immediately releasing the same report queue or migration.
Choose the repair that matches the evidence:
- A single expensive query: cancel it if still active, inspect its plan in a safe environment, add or correct indexes, reduce the result set, or schedule it away from peak load.
- Oversized global
work_mem: lower the global value and grant a higher session- or role-specific value only to tested analytical work that needs it. - Excessive concurrency: reduce web and worker pool totals, cap job parallelism, and keep administrative and backup headroom.
- Parallel-query multiplication: test a lower
max_parallel_workers_per_gatherfor the affected role or workload rather than disabling parallelism blindly across the cluster. - Maintenance collision: avoid overlapping large index builds, restores, backups, and peak application load; review maintenance memory and worker counts.
- Container limit below measured need: raise it only if the host has durable headroom for the database plus every neighbouring service.
- VPS genuinely too small: move the database or workload to a host with adequate memory instead of hiding sustained pressure with repeated restarts.
Some PostgreSQL settings reload; others require restart. Check pg_settings.context and pending_restart, make one controlled change, and let Coolify recreate or restart the resource only when required.
Do not “fix” OOM by disabling the OOM killer
Docker supports controls related to OOM behaviour, but protecting an unbounded database process can move the failure to Docker, Coolify, SSH, or another critical host process. PostgreSQL documentation also discusses Linux overcommit and OOM score adjustment, but these are host-design decisions—not first-response toggles for an undersized shared VPS.
Likewise, swap can provide a short buffer but is not substitute RAM for a busy database. Heavy swapping can turn a sharp failure into severe latency, health-check failures, and timeouts. Measure it and size the system for the workload.
6. Set Coolify limits from measured headroom
In the PostgreSQL resource, open Configuration → Resource Limits. Coolify supports a soft memory reservation, maximum memory limit, swap limit, swappiness, and CPU controls. A value of zero means no configured limit for the relevant limit fields.
A useful hard limit must satisfy two constraints:
database limit > tested peak database demand + operational burst margin
database limit + other services + builds/backups + host reserve
< safely usable host memory
Do not select a number from a generic tuning chart. Measure the actual application mix, including startup recovery, migrations, backups, autovacuum, reporting, and a representative concurrency peak. Save the Coolify limit, restart or redeploy the database as required, and confirm the running container received the intended values.
7. Verify more than a green health badge
- Confirm the database container starts once and its restart count remains stable.
- Confirm Docker no longer reports an OOM kill and the kernel records no new OOM event.
- Wait for PostgreSQL crash recovery, if any, to complete before judging readiness.
- Run a database-backed application read and one controlled write with read-back.
- Run the worker or report path that previously triggered the incident at bounded concurrency.
- Watch Coolify Metrics and host memory through the full operation, not only its first seconds.
- Confirm connection counts remain within the planned pool budget.
- Run a fresh scheduled backup and verify completion.
- Check that application latency, temporary-file I/O, and swap activity remain acceptable.
A passing PostgreSQL health check proves that its configured probe succeeded. It does not prove that a representative query fits the memory budget, a backup can complete, or the application can read and write safely.
Compact incident order
- Record the failure time, container state, restart count, logs, metrics, and Coolify limits.
- Prove container OOM, host OOM, or a different exit cause.
- Preserve the persistent volume and stop the triggering workload before recovery.
- Read effective PostgreSQL memory, connection, maintenance, and parallel-worker settings.
- Match the spike to queries, jobs, migrations, maintenance, backups, or a deployment.
- Reduce the narrow multiplier: query cost, concurrency, pool size, parallelism, or an oversized setting.
- Raise the Coolify limit or host capacity only from measured demand and real headroom.
- Restart once when required, then verify reads, writes, jobs, backups, memory, and restart stability.
The durable fix is not “give PostgreSQL all the RAM.” It is to make the memory boundary and workload multipliers explicit. Coolify controls the container boundary; PostgreSQL controls shared and per-operation memory; the application controls concurrency; and every service still shares one host. When those budgets agree, an analytical query or routine backup is far less likely to become a database outage.
Authoritative references: Coolify documents database resource limits, database health checks, and its standalone database model. Docker explains container memory constraints and OOM risk. PostgreSQL documents shared, query, maintenance, and parallel-worker memory, Linux memory overcommit and OOM behaviour, pg_stat_activity, and backend memory-context logging.