Skip to main content

Read Logs and Traces

The gNB planes (CU-CP, CU-UP and every DU), the controller, the RANN-P forwarder and CU-IP send their logs and traces over OTLP to one collector in the monitoring namespace. Logs land in ClickHouse, traces in Tempo, and Grafana reads both. The names and attributes every span carries, and the topology in detail, are the span schema; this page is how you read them.

Not every component reaches the collector:

ComponentLogsTraces
CU-CP, CU-UP, each DU (the gnb image)ClickHouse, and the CU-CP's own file /tmp/cu_cp.logTempo
racora-controllerClickHouseTempo
rann-p-forwarder (CU-CP sidecar)ClickHouseTempo
cuipClickHouse: its library defaults to the in-cluster collector when no endpoint is setTempo
racora-core-controllerkubectl logs only: it carries no OpenTelemetrynone
the device pluginskubectl logs onlynone

Both stores keep 24 hours (racora-monitoring.clickhouse.logsTtl, tempo.blockRetention) and default to emptyDir, so a pod restart clears them. Set racora-monitoring.clickhouse.storage.persistent and tempo.storage.persistent to keep them on a PersistentVolumeClaim (chart values).

Is the Stack Up​

kubectl get pods -n monitoring # clickhouse, tempo, otel-collector, grafana: 1/1 Running
curl -s http://<node-ip>:30300/api/health

Once the collector is up, the Failed to export logs to otel-collector… lines in every producer's own log stop.

Live Logs, per Component​

kubectl logs -n racora-system deploy/racora-controller -f # reconcile, mobility sync, decision loop
kubectl logs -n racora-system deploy/racora-core-controller -f # subscriber provisioning
kubectl exec -n centralized-unit deploy/cu-cp -c cu-cp -- tail -F /tmp/cu_cp.log # the CU-CP's own log file
kubectl logs -n centralized-unit deploy/cuip -f # substrate, engines, decisions
kubectl logs -n distributed-unit deploy/du-<cell name> -f # one DU per cell

The CU-CP log file is where every runtime command reports its effect (Added neighbor relation…, Updated report config id=…, Set periodic report…), where handovers narrate themselves, and where the RNTI of an attached UE is printed (Updated UE with …). The controller's overlay sets log.force_flush, so the file is flushed on every write.

Durable Logs: ClickHouse​

The clickhouse-client is in the ClickHouse pod:

kubectl exec -n monitoring deploy/clickhouse -- clickhouse-client -q "
SELECT Timestamp, ServiceName, Body
FROM otel.otel_logs
WHERE ServiceName = 'cu-cp' AND Body LIKE '%Handover%'
ORDER BY Timestamp DESC LIMIT 50 FORMAT PrettyCompact"

ServiceName is one of cu-cp, cu-up, du, racora-controller, rann-p-forwarder, cuip. Markers worth knowing:

You want to seeLook for (ServiceName='cu-cp')
a handover, start to finishTrigger intra-CU (inter-DU) handover from source_du=A to target_du=B, then "Intra CU Handover Routine" finished successfully, then "Intra CU Handover Target Routine" finished successfully
a handover refused because the target is lockedIgnoring Handover Request. Cause: Target cell with pci=N is administratively deactivated
a mobility change appliedAdded neighbor relation…, Updated report config id=…, Set periodic report…
why a runtime command was rejectedthe line right before the WS error; the WS reply is deliberately coarse
the periodic RRC/NGAP countersthe periodic metrics lines (period racora-controller.cucpMetricsPeriodMs, default 5 s) that the dashboard counts

Grafana's Explore with the ClickHouse datasource runs the same SQL interactively.

Dashboards​

Grafana is http://<node-ip>:30300 (a NodePort; racora-monitoring.grafana.service.type switches it to ClusterIP for a port-forward). Anyone who reaches it reads dashboards as the anonymous Viewer (racora-monitoring.grafana.anonymousAccess); the admin login is Grafana's default, admin with password admin, and the chart offers no way to set another that survives a pod restart, so beyond a lab keep the Service ClusterIP and port-forward. Dashboards, folder Racora, dashboard RAN Mobility: handovers completed, reconfiguration timeouts, misrouted handovers, triggers, A3 decisions and measurement reports, and the recent mobility events, all counted from the CU-CP log markers above. Log in as admin to edit or add dashboards.

Two things the chart does not persist: Grafana's data directory is an emptyDir, so a password or dashboard changed in the UI reverts when the pod restarts (provisioned dashboards return; this is why the default password stands), and the ClickHouse datasource is a plugin Grafana installs at start from grafana.com (GF_INSTALL_PLUGINS), so the pod needs internet egress; without it the log dashboards have no datasource (offline installs).

Traces: Tempo​

Grafana, then Explore, datasource Tempo: search by service (racora-controller, cu-cp) or by span name. The two shapes you will look at most:

  • An operator edit of spec.neighbors or spec.mobility: cucp.mobility_sync (controller), then one cucp.mobility_sync.cmd per command, then the CU-CP's cucp.<cmd>.execute, joined by the traceparent the controller puts in every command.
  • An autonomous decision: CU-IP's rannd.emit, then the controller's dispatch.consider and, for a PCI retune, cucp.cell_lock, cucp.drain, dispatch.apply, dispatch.rollout_wait and cucp.cell_unlock.

A failed operation has span status ERROR; a tolerated or skipped one does not, and says why in its outcome attribute (cmd.outcome, dispatch.outcome, mobility.sync_outcome).

From a Trace to Its Logs​

Every log line written inside an active span carries that span's identity, so a trace id from Tempo selects the logs of the whole operation:

kubectl exec -n monitoring deploy/clickhouse -- clickhouse-client -q "
SELECT Timestamp, ServiceName, Body FROM otel.otel_logs
WHERE TraceId = '<32-hex trace id>' ORDER BY Timestamp FORMAT PrettyCompact"

This join is manual by design: Grafana's built-in span-to-logs link does not support the ClickHouse datasource. To trace a hand-sent runtime command, add a top-level traceparent (00-<32 hex>-<16 hex>-01) next to cmd; the CU-CP's span lands under that id.