Skip to content

Monitoring ​

A Mankomail instance tells you how it is doing through two HTTP endpoints and its logs. This page describes both, what is worth watching, how to size the database, and what the instance does on its own after a restart or an outage.

Health endpoints ​

Every application process, whatever its APP_ROLE, answers two endpoints on its HTTP port. Neither requires authentication.

EndpointQuestion it answersTouches the database
/healthzIs the process alive?No
/readyzCan the process serve requests?Yes

/healthz always answers 200 while the process runs:

json
{ "status": "ok", "brand": "Mankomail", "version": "1.4.0", "uptimeSeconds": 5321.4 }

brand is the value of BRAND_NAME, version the value of APP_VERSION, and uptimeSeconds the time since the process started.

/readyz runs two checks, in order: the database answers a query, then every migration of the image has been applied. It answers 200 when both pass:

json
{
  "status": "ready",
  "checks": [
    { "name": "database", "ok": true },
    { "name": "migrations", "ok": true }
  ]
}

and 503 otherwise, with the reason in detail:

json
{
  "status": "not_ready",
  "checks": [
    { "name": "database", "ok": false, "detail": "connect ECONNREFUSED 172.18.0.2:5432" },
    { "name": "migrations", "ok": false, "detail": "database unreachable" }
  ]
}

When the database is reachable but migrations are missing, the migrations check carries pending: followed by the file names of the missing migrations.

Which endpoint to use where ​

  • Liveness probe (Docker health check, Kubernetes livenessProbe, process supervisor): /healthz. A probe that touched the database would restart healthy processes in a loop during a short database outage.
  • Readiness probe and load balancer (Kubernetes readinessProbe, upstream checks of a proxy): /readyz.
  • External availability check: <PUBLIC_BASE_URL>/readyz, from outside the server, so that the reverse proxy and the certificate are checked too.

The image declares a Docker HEALTHCHECK on /healthz: every 30 seconds, 5-second timeout, unhealthy after 3 failures, with a 20-second start period. docker compose ps shows the result.

The detail field of /readyz can contain database error messages and migration file names. If you prefer not to expose them publicly, restrict /readyz to your monitoring addresses at the reverse proxy, and probe it locally on 127.0.0.1:<APP_PORT>.

Logs ​

The application writes its logs on the standard output, one JSON object per line. With the reference Compose file, read them with:

bash
docker compose -f compose.reference.yaml logs -f app

Each line carries:

FieldContent
levelfatal, error, warn, info, debug or trace
timeISO 8601 timestamp
roleThe APP_ROLE of the process
versionThe APP_VERSION of the process
msgThe message
moduleThe part of the application that wrote the line, when it has one

HTTP requests are logged at the info level, with a request identifier (reqId) shared by the request line and the response line.

  • LOG_LEVEL sets the minimum level: fatal, error, warn, info (default), debug, trace or silent.
  • LOG_FORMAT accepts json (one JSON object per line, for log collectors) and pretty (readable lines, for development).
  • Secrets and personal data are redacted: authorization headers, cookies, passwords, tokens, API keys, and the subject, body and recipients of emails are replaced with [redacted].
  • A fatal error at startup, such as a failed migration or an unreachable database, is written as plain text starting with fatal: failed to start, followed by the error.

To filter the JSON lines, use jq:

bash
docker compose -f compose.reference.yaml logs --no-log-prefix app \
  | jq -c 'select(.level == "error" or .level == "fatal")'

The reference Compose file keeps at most five files of 10 MB of logs per container. To keep them longer, ship them to a log system (Loki, Elasticsearch, a syslog server) with a Docker logging driver or a collector.

Log lines worth knowing ​

Log lineLevelWhat it means
configuration: …warnA variable has an invalid value and its default is used instead.
ENCRYPTION_KEY is not set / ENCRYPTION_KEY is invalidwarnCredential storage is disabled. See The encryption key.
database schema is up to dateinfoMigrations are applied.
postgres is not accepting connections yet, retryingwarnThe database is not ready yet at boot. Normal for a few seconds.
expired job leases reclaimed (a worker died holding them)warnJobs held by a process that died were handed back to the queue. Expected after a crash or a forced stop.
claim failed, scheduler tick failederrorThe background loops cannot reach the database. Repeated every second during a database outage.
postgres pool: connection lost while in usewarnA database connection was cut during a query.
instance maintenance doneinfoThe hourly maintenance ran.
APP_ROLE=api: this process runs no workerwarnThis process runs no background work: make sure a worker process runs somewhere.

A few error lines during a database restart are expected. The same lines repeated for minutes, while PostgreSQL is up, are not.

Metrics ​

The current version exposes no metrics endpoint: there is no Prometheus or OpenTelemetry export. Watch the instance through the health endpoints, the logs, and a few measures you collect yourself.

MeasureHow to get itWhy
/readyz statusExternal HTTP checkAvailability of the instance.
Free disk spaceYour system monitoringThe object storage and PostgreSQL grow with the mailboxes. A full disk stops everything.
PostgreSQL sizeSELECT pg_size_pretty(pg_database_size(current_database()));Drives the database sizing below.
Age of the oldest ready jobSQL query belowThe best sign that background work is stuck.
Jobs that exhausted their attemptsSQL query belowWork that needs a human look.

The job queue lives in the jobs table. This query gives its state:

sql
SELECT
  count(*) FILTER (WHERE dead_at IS NULL AND claimed_by IS NULL AND run_at <= now()) AS ready,
  count(*) FILTER (WHERE dead_at IS NULL AND claimed_by IS NOT NULL)                 AS running,
  count(*) FILTER (WHERE dead_at IS NOT NULL)                                        AS dead,
  now() - min(run_at) FILTER (WHERE dead_at IS NULL AND claimed_by IS NULL AND run_at <= now())
    AS oldest_ready_age
FROM jobs;

Run it with:

bash
docker compose -f compose.reference.yaml exec postgres \
  sh -c 'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB"'
  • oldest_ready_age stays at a few seconds on a healthy instance. Several minutes means no process is taking work: check that a worker or all process runs and that its logs show no claim failed.
  • dead counts jobs that exhausted their attempts. They are kept 30 days, then deleted by the maintenance. A number that keeps growing deserves a look at their kind and last_error.

The database schema is internal and may change from one version to the next: check these queries after an upgrade.

Workflow failures are not an operations matter: members see them on the Activity page and in their notifications. See Errors and retries.

Sizing ​

Load tests of the ingestion path show that a single application process handles far more incoming email than an organisation of a hundred mailboxes receives: the background workers are not the first limit. PostgreSQL's memory is.

  • A message costs about 1.7 KB in PostgreSQL, indexes included. Email bodies and attachments are in the object storage, which is by far the largest consumer of disk.
  • The PostgreSQL image starts with shared_buffers=128MB, whatever memory the container has. Around 50,000 messages, the working set no longer fits in that cache, and reading older threads starts to hit the disk.
Messages in the databasePostgreSQL sizeshared_buffersPostgreSQL container memory
50,000~110 MB256 MB2 GB (default)
500,000~1.1 GB1 GB4 GB
5,000,000~11 GB4 GB16 GB

Set the PostgreSQL parameters in your Compose file, and raise POSTGRES_MEMORY to match:

yaml
services:
  postgres:
    command: >
      postgres -c shared_buffers=1GB -c effective_cache_size=3GB
               -c work_mem=16MB -c max_wal_size=4GB

A larger max_wal_size spaces out forced checkpoints, which cause occasional slow writes during a large initial synchronisation.

InstanceMailboxesSuggested set-up
Evaluation1 to 5The reference Compose file as is.
Team10 to 50shared_buffers=512MB, max_wal_size=4GB, PostgreSQL container with 4 GB. Daily backups.
Large100 and morePostgreSQL outside the Compose file (managed service or dedicated machine), separate api and worker processes, external object storage. Keep DATABASE_POOL_MAX × number of processes below max_connections.

DATABASE_STATEMENT_TIMEOUT_MS (30 seconds by default) cancels any query that runs longer, so that a runaway query cannot hold a connection forever. Migrations and backups are not affected.

Restarts and recovery ​

The instance is built to be stopped and killed. All of its state lives in PostgreSQL, so a restart loses no work.

Clean stop. On SIGTERM or SIGINT, the process stops accepting requests, stops its scheduler, lets the jobs it holds finish for up to 30 seconds, closes its database connections, and logs shutdown complete. Give the container at least 40 seconds to stop: see Work in progress during an upgrade.

Crash or forced stop. Every job a process holds has a 60-second lease, renewed while the job runs. When a process dies, its leases expire, and a process that runs background work hands the jobs back to the queue (expired job leases reclaimed). A job that keeps failing ends among the dead jobs instead of killing processes one after the other.

Boot before the database. At boot, the application waits up to 60 seconds for PostgreSQL, retrying with a growing delay. A wrong password, an unknown role or a missing database stop the boot at once, with the server's message: waiting would not fix them. After 60 seconds without a database, the process stops with the last PostgreSQL error, and Docker restarts it.

Database outage while running. A PostgreSQL restart, crash or network cut does not stop the application:

During the outageWhen the database is back
/healthz200200
/readyz503, with the PostgreSQL error in detail200 at the first accepted connection, without restarting the application
HTTP requests that need the databaseFail with an errorWork again
Background workclaim failed and scheduler tick failed, every secondResumes on its own; interrupted jobs are handed back to the queue when their lease expires

If /readyz stays at 503 while PostgreSQL accepts connections, read its detail: it usually points at an authentication or migration problem. An application container that restarts after each database outage is not expected behaviour.