> ## Documentation Index
> Fetch the complete documentation index at: https://docs.tofunmiadewuyi.com/custos/llms.txt
> Use this file to discover all available pages before exploring further.

# Observability

> Structured logs, OpenTelemetry traces and metrics, and the daemon telemetry relay.

## Logs

The control plane and the daemon both write structured JSON logs to stderr. `CUSTOS_LOG_LEVEL`
selects `debug`, `info`, `warn`, or `error`; the default is `info`.

Logs emitted inside a traced operation carry `trace_id` and `span_id`, so a request can be followed
from a log record into its trace.

## Traces and metrics

Both binaries create OpenTelemetry traces for HTTP requests, PostgreSQL calls, outbound HTTP calls,
daemon authentication, and discrete WebSocket message operations. A small set of application metrics
is aggregated in memory.

On the control plane, OTLP export is **disabled** until you set an endpoint:

```bash theme={null}
export OTEL_EXPORTER_OTLP_ENDPOINT=https://telemetry.example.com
export OTEL_EXPORTER_OTLP_HEADERS='authorization=Bearer <token>'
```

The general endpoint enables logs, traces, and metrics over OTLP HTTP/protobuf. Signal-specific
variables are available when you want only some signals — see
[Configuration](/custos/custos/control-plane/configuration).

`OTEL_METRIC_EXPORT_INTERVAL` overrides the default 60-second metric export interval, in
milliseconds, for both the control plane and the daemon.

## How daemon telemetry gets out

Daemons never connect to an OTLP backend and open no additional connection. They batch standard OTLP
protobufs and forward only the signals the control plane has enabled, over the existing
authenticated WebSocket. The control plane relays those batches to the same signal-specific
destinations it uses for its own telemetry.

Daemon telemetry carries the enrolled `host.id` as a resource attribute, so fleet-wide queries can
group by host.

Design properties worth knowing when you are debugging a quiet fleet:

* Core messages and telemetry sit in **separate bounded queues**. The writer gives core messages
  priority and **drops telemetry** rather than blocking Custos when a telemetry queue fills.
* Ping/pong is independent of both, so keepalives are never starved by telemetry volume.
* Telemetry is best-effort and never persists to the daemon's state directory.

In other words: missing daemon telemetry under load is expected behaviour, not a fault. Access
control and secret delivery are never delayed for a metric.

## Grafana

A reusable dashboard and its import instructions ship in the repository:

* [`observability/grafana/dashboards/custos-overview.json`](https://github.com/tofunmiadewuyi/custos/blob/main/observability/grafana/dashboards/custos-overview.json)
* [`observability/README.md`](https://github.com/tofunmiadewuyi/custos/blob/main/observability/README.md)

## A note on collector placement

When the control plane runs directly on the host alongside a collector, `127.0.0.1:4318` reaches it.
A containerized control plane needs the collector's service name on the shared telemetry network
instead, for example `http://alloy:4318`.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.