Monitoring
The API emits redacted structured logs and aggregate metrics. Monitoring does not send family data to a telemetry provider. Keep the collector, storage and alert destination within your approved hosting boundary.
Protected metrics
Set METRICS_TOKEN to 32 random bytes encoded as lowercase hex:
openssl rand -hex 32Save it in the API's secret environment and in the collector's credential file. Do not use an account session or a personal API token. With no METRICS_TOKEN, GET /metrics returns 404. A wrong or missing Authorization: Bearer … header returns 401; query-string tokens are never accepted. Responses have Cache-Control: no-store. Restrict /metrics to the private monitoring network at your reverse proxy as an additional control.
The output is Prometheus text format with fixed labels: status classes, operation names, queue states, allocation types and collection names. It contains no account/family IDs, URLs, object keys, request content, tokens or exception text.
- Counters reset on API restart. Use
rate()orincrease(), not subtraction across restarts. Duration histograms use seconds and cumulative buckets. /health,/readyand/metricsdo not contribute to request rates, so scrape traffic does not hide errors during quiet periods.fellesly_readychecks the same database/startup conditions as/ready. A probe exception produces zero.fellesly_writer_healthyonly reports the persistence fence. Neither probes email, storage, push or payment providers.- Queued records are gauges, including terminal states: retention can reduce them. Due age measures delay since the next scheduled attempt, not original creation.
- A job's last successful run does not establish that every provider operation succeeded. Partial failures still have separate redacted logs.
Optional local collector
The repository includes deploy/docker-compose.monitoring.yml, which adds Prometheus, Alertmanager and bounded local Docker logs for the API. It must be merged with a base Compose file. It does not alter the single-writer API model.
Set these additional deployment values:
| Variable | Value |
|---|---|
METRICS_TOKEN | The API's generated monitoring secret |
METRICS_TOKEN_FILE | Absolute host path to a file containing the same secret |
ALERTMANAGER_CONFIG_FILE | Absolute host path to your Alertmanager receiver configuration |
Keep both files outside the repository, in a host directory accessible only to its owner. The container users must be able to read the mounted files. For example, a root-owned directory with mode 0700 and contained files with mode 0444 allows Docker to mount readable files while preventing other host users from traversing the directory. Do not make the parent directory world-readable.
Configure the receiver in the protected Alertmanager file, for example:
route:
receiver: operations
group_by: [alertname, job, instance]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
receivers:
- name: operations
webhook_configs:
- url: https://your-approved-incident-receiver.example/alerts
send_resolved: trueReplace that example URL before starting the collector. Treat a secret webhook URL as a credential. You may instead configure another supported Alertmanager receiver; any referenced credential files also need read-only mounts. No default receiver that silently discards alerts is installed.
For a self-hosted Compose deployment:
docker compose --env-file deploy/.env \
-f deploy/docker-compose.yml -f deploy/docker-compose.monitoring.yml \
up -d --buildFor the GHCR/Coolify stack, merge the monitoring overlay with deploy/docker-compose.coolify.yml, keeping the existing project's name and volumes. Ensure the configuration files exist on the deployment host before deploying. Changing the project name can create a separate database volume; do not launch a second API against the existing database.
Prometheus and Alertmanager bind only to host loopback ports 9090 and 9093. Use an SSH tunnel to inspect them; do not add public Coolify domains to these services. The API scrape uses the private Compose network. Use HTTPS if the collector is on a different host. Prometheus retains at most 14 days or 1 GB of samples, whichever limit is reached first. API logs rotate at three files of 10 MB. These bounds are operational defaults, not a retention policy for family data or backups.
Alerts and validation
The supplied rules cover missing metrics, failed readiness, failed operations, delayed push attempts, orphan cleanup failures and elevated server errors with a minimum traffic floor. Review thresholds against your staging traffic. Add a memory threshold based on your actual API container limit and independent off-host uptime and missed-backup checks. A monitor on the API's host cannot notify you when the whole host is down.
Validate configuration before deployment:
promtool check config deploy/monitoring/prometheus.yml
promtool check rules deploy/monitoring/alerts.yml
promtool test rules deploy/monitoring/alerts.test.yml
amtool check-config /absolute/path/to/alertmanager.ymlThe Prometheus config uses container paths for the rules and credential file. Run its config check inside the container, or copy the config and replace those paths for local validation. The rule tests run directly from the repository.
On staging, verify that Prometheus reports the fellesly target as UP, stop the API for more than three minutes and confirm both alert delivery and subsequent resolution. Repeat with a failing readiness probe and a delayed queue using isolated test data. Verify the receiver's own availability and review its data handling. Local checks and unit tests do not establish real incident delivery or production recovery readiness.