English
Monitoring
A Mankomail instance tells you how it is doing through two HTTP endpoints and its logs. This page describes both, what is worth watching, how to size the database, and what the instance does on its own after a restart or an outage.
Health endpoints
Every application process, whatever its APP_ROLE, answers two endpoints on its HTTP port. Neither requires authentication.
| Endpoint | Question it answers | Touches the database |
|---|---|---|
/healthz | Is the process alive? | No |
/readyz | Can the process serve requests? | Yes |
/healthz always answers 200 while the process runs:
json
{ "status": "ok", "brand": "Mankomail", "version": "1.4.0", "uptimeSeconds": 5321.4 }brand is the value of BRAND_NAME, version the value of APP_VERSION, and uptimeSeconds the time since the process started.
/readyz runs two checks, in order: the database answers a query, then every migration of the image has been applied. It answers 200 when both pass:
json
{
"status": "ready",
"checks": [
{ "name": "database", "ok": true },
{ "name": "migrations", "ok": true }
]
}and 503 otherwise, with the reason in detail:
json
{
"status": "not_ready",
"checks": [
{ "name": "database", "ok": false, "detail": "connect ECONNREFUSED 172.18.0.2:5432" },
{ "name": "migrations", "ok": false, "detail": "database unreachable" }
]
}When the database is reachable but migrations are missing, the migrations check carries pending: followed by the file names of the missing migrations.
Which endpoint to use where
- Liveness probe (Docker health check, Kubernetes
livenessProbe, process supervisor):/healthz. A probe that touched the database would restart healthy processes in a loop during a short database outage. - Readiness probe and load balancer (Kubernetes
readinessProbe, upstream checks of a proxy):/readyz. - External availability check:
<PUBLIC_BASE_URL>/readyz, from outside the server, so that the reverse proxy and the certificate are checked too.
The image declares a Docker HEALTHCHECK on /healthz: every 30 seconds, 5-second timeout, unhealthy after 3 failures, with a 20-second start period. docker compose ps shows the result.
The detail field of /readyz can contain database error messages and migration file names. If you prefer not to expose them publicly, restrict /readyz to your monitoring addresses at the reverse proxy, and probe it locally on 127.0.0.1:<APP_PORT>.
Logs
The application writes its logs on the standard output, one JSON object per line. With the reference Compose file, read them with:
bash
docker compose -f compose.reference.yaml logs -f appEach line carries:
| Field | Content |
|---|---|
level | fatal, error, warn, info, debug or trace |
time | ISO 8601 timestamp |
role | The APP_ROLE of the process |
version | The APP_VERSION of the process |
msg | The message |
module | The part of the application that wrote the line, when it has one |
HTTP requests are logged at the info level, with a request identifier (reqId) shared by the request line and the response line.
LOG_LEVELsets the minimum level:fatal,error,warn,info(default),debug,traceorsilent.LOG_FORMATacceptsjson(one JSON object per line, for log collectors) andpretty(readable lines, for development).- Secrets and personal data are redacted: authorization headers, cookies, passwords, tokens, API keys, and the subject, body and recipients of emails are replaced with
[redacted]. - A fatal error at startup, such as a failed migration or an unreachable database, is written as plain text starting with
fatal: failed to start, followed by the error.
To filter the JSON lines, use jq:
bash
docker compose -f compose.reference.yaml logs --no-log-prefix app \
| jq -c 'select(.level == "error" or .level == "fatal")'The reference Compose file keeps at most five files of 10 MB of logs per container. To keep them longer, ship them to a log system (Loki, Elasticsearch, a syslog server) with a Docker logging driver or a collector.
Log lines worth knowing
| Log line | Level | What it means |
|---|---|---|
configuration: … | warn | A variable has an invalid value and its default is used instead. |
ENCRYPTION_KEY is not set / ENCRYPTION_KEY is invalid | warn | Credential storage is disabled. See The encryption key. |
database schema is up to date | info | Migrations are applied. |
postgres is not accepting connections yet, retrying | warn | The database is not ready yet at boot. Normal for a few seconds. |
expired job leases reclaimed (a worker died holding them) | warn | Jobs held by a process that died were handed back to the queue. Expected after a crash or a forced stop. |
claim failed, scheduler tick failed | error | The background loops cannot reach the database. Repeated every second during a database outage. |
postgres pool: connection lost while in use | warn | A database connection was cut during a query. |
instance maintenance done | info | The hourly maintenance ran. |
APP_ROLE=api: this process runs no worker | warn | This process runs no background work: make sure a worker process runs somewhere. |
A few error lines during a database restart are expected. The same lines repeated for minutes, while PostgreSQL is up, are not.
Metrics
The current version exposes no metrics endpoint: there is no Prometheus or OpenTelemetry export. Watch the instance through the health endpoints, the logs, and a few measures you collect yourself.
| Measure | How to get it | Why |
|---|---|---|
/readyz status | External HTTP check | Availability of the instance. |
| Free disk space | Your system monitoring | The object storage and PostgreSQL grow with the mailboxes. A full disk stops everything. |
| PostgreSQL size | SELECT pg_size_pretty(pg_database_size(current_database())); | Drives the database sizing below. |
| Age of the oldest ready job | SQL query below | The best sign that background work is stuck. |
| Jobs that exhausted their attempts | SQL query below | Work that needs a human look. |
The job queue lives in the jobs table. This query gives its state:
sql
SELECT
count(*) FILTER (WHERE dead_at IS NULL AND claimed_by IS NULL AND run_at <= now()) AS ready,
count(*) FILTER (WHERE dead_at IS NULL AND claimed_by IS NOT NULL) AS running,
count(*) FILTER (WHERE dead_at IS NOT NULL) AS dead,
now() - min(run_at) FILTER (WHERE dead_at IS NULL AND claimed_by IS NULL AND run_at <= now())
AS oldest_ready_age
FROM jobs;Run it with:
bash
docker compose -f compose.reference.yaml exec postgres \
sh -c 'psql -U "$POSTGRES_USER" -d "$POSTGRES_DB"'oldest_ready_agestays at a few seconds on a healthy instance. Several minutes means no process is taking work: check that aworkerorallprocess runs and that its logs show noclaim failed.deadcounts jobs that exhausted their attempts. They are kept 30 days, then deleted by the maintenance. A number that keeps growing deserves a look at theirkindandlast_error.
The database schema is internal and may change from one version to the next: check these queries after an upgrade.
Workflow failures are not an operations matter: members see them on the Activity page and in their notifications. See Errors and retries.
Sizing
Load tests of the ingestion path show that a single application process handles far more incoming email than an organisation of a hundred mailboxes receives: the background workers are not the first limit. PostgreSQL's memory is.
- A message costs about 1.7 KB in PostgreSQL, indexes included. Email bodies and attachments are in the object storage, which is by far the largest consumer of disk.
- The PostgreSQL image starts with
shared_buffers=128MB, whatever memory the container has. Around 50,000 messages, the working set no longer fits in that cache, and reading older threads starts to hit the disk.
| Messages in the database | PostgreSQL size | shared_buffers | PostgreSQL container memory |
|---|---|---|---|
| 50,000 | ~110 MB | 256 MB | 2 GB (default) |
| 500,000 | ~1.1 GB | 1 GB | 4 GB |
| 5,000,000 | ~11 GB | 4 GB | 16 GB |
Set the PostgreSQL parameters in your Compose file, and raise POSTGRES_MEMORY to match:
yaml
services:
postgres:
command: >
postgres -c shared_buffers=1GB -c effective_cache_size=3GB
-c work_mem=16MB -c max_wal_size=4GBA larger max_wal_size spaces out forced checkpoints, which cause occasional slow writes during a large initial synchronisation.
| Instance | Mailboxes | Suggested set-up |
|---|---|---|
| Evaluation | 1 to 5 | The reference Compose file as is. |
| Team | 10 to 50 | shared_buffers=512MB, max_wal_size=4GB, PostgreSQL container with 4 GB. Daily backups. |
| Large | 100 and more | PostgreSQL outside the Compose file (managed service or dedicated machine), separate api and worker processes, external object storage. Keep DATABASE_POOL_MAX × number of processes below max_connections. |
DATABASE_STATEMENT_TIMEOUT_MS (30 seconds by default) cancels any query that runs longer, so that a runaway query cannot hold a connection forever. Migrations and backups are not affected.
Restarts and recovery
The instance is built to be stopped and killed. All of its state lives in PostgreSQL, so a restart loses no work.
Clean stop. On SIGTERM or SIGINT, the process stops accepting requests, stops its scheduler, lets the jobs it holds finish for up to 30 seconds, closes its database connections, and logs shutdown complete. Give the container at least 40 seconds to stop: see Work in progress during an upgrade.
Crash or forced stop. Every job a process holds has a 60-second lease, renewed while the job runs. When a process dies, its leases expire, and a process that runs background work hands the jobs back to the queue (expired job leases reclaimed). A job that keeps failing ends among the dead jobs instead of killing processes one after the other.
Boot before the database. At boot, the application waits up to 60 seconds for PostgreSQL, retrying with a growing delay. A wrong password, an unknown role or a missing database stop the boot at once, with the server's message: waiting would not fix them. After 60 seconds without a database, the process stops with the last PostgreSQL error, and Docker restarts it.
Database outage while running. A PostgreSQL restart, crash or network cut does not stop the application:
| During the outage | When the database is back | |
|---|---|---|
/healthz | 200 | 200 |
/readyz | 503, with the PostgreSQL error in detail | 200 at the first accepted connection, without restarting the application |
| HTTP requests that need the database | Fail with an error | Work again |
| Background work | claim failed and scheduler tick failed, every second | Resumes on its own; interrupted jobs are handed back to the queue when their lease expires |
If /readyz stays at 503 while PostgreSQL accepts connections, read its detail: it usually points at an authentication or migration problem. An application container that restarts after each database outage is not expected behaviour.