All checks were successful
CI / lint (push) Successful in 33s
CI / test (push) Successful in 2m6s
Deploy / notebooks (push) Has been skipped
CI / notebooks-smoke (push) Successful in 1m31s
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / llm (push) Successful in 1m21s
Deploy / mc (push) Has been skipped
Deploy / api (push) Successful in 1m46s
Infra CI / docs (push) Successful in 21s
Infra CI / llm (push) Successful in 15s
Infra CI / mc (push) Successful in 14s
Deploy / report (push) Successful in 11s
Infra CI / zotero (push) Successful in 17s
Infra CI / notebooks (push) Successful in 50s
Infra CI / api (push) Successful in 19s
The P26 telemetry path never reached Prometheus: setup_meter_provider only installs an in-process PrometheusMetricReader, nothing served the registry (api and llm answered 404 on /metrics), the images never installed the perf extra, STACK_TELEMETRY was off, and no scrape target existed — so the data-pipelines request-rate panel was empty from the day it was written. Now: perf.middleware.instrument mounts GET /metrics (prometheus_client registry), the api and llm images install --extra perf, compose sets STACK_TELEMETRY=true on both, services.yml scrapes api:8000 and llm:8000, and both dashboards' request-rate panels query http_server_duration_milliseconds_count (what the FastAPI instrumentor emits; stack_http_server_requests_total never existed). Verified live: stack_llm_dispatch_total is queryable in Prometheus with job=llm.
61 lines
1.4 KiB
YAML
61 lines
1.4 KiB
YAML
# ── Scrape Target Registry ───────────────────────────
|
|
# To add a service: append one entry below.
|
|
# Pattern: container_name:port, job label, optional __metrics_path__.
|
|
|
|
# Traefik
|
|
- targets: ['traefik:8080']
|
|
labels:
|
|
job: traefik
|
|
__metrics_path__: /metrics
|
|
|
|
# Tempo
|
|
- targets: ['tempo:3200']
|
|
labels:
|
|
job: tempo
|
|
__metrics_path__: /metrics
|
|
|
|
# Loki
|
|
- targets: ['loki:3100']
|
|
labels:
|
|
job: loki
|
|
|
|
# NVIDIA GPU (DCGM)
|
|
- targets: ['nvidia-exporter:9400']
|
|
labels:
|
|
job: nvidia-gpu
|
|
|
|
# Nessie
|
|
- targets: ['nessie:9000']
|
|
labels:
|
|
job: nessie
|
|
__metrics_path__: /q/metrics
|
|
|
|
# Polaris
|
|
- targets: ['polaris:8182']
|
|
labels:
|
|
job: polaris
|
|
__metrics_path__: /q/metrics
|
|
|
|
# Trino — /metrics requires auth (Trino rejects unauthenticated scrapes
|
|
# even with no auth configured). Skipped until the openmetrics catalog
|
|
# is wired with a shared secret. Loki still captures query lifecycle
|
|
# events for the data-lake dashboard.
|
|
|
|
# OTel Collector
|
|
- targets: ['otel-collector:8889']
|
|
labels:
|
|
job: otel-collector
|
|
|
|
# Stack API + LLM service — perf.middleware serves the OTel Prometheus
|
|
# registry at /metrics (#579): http_server_* from the FastAPI
|
|
# instrumentor, stack_llm_* dispatch/embedding/chat/index instruments.
|
|
- targets: ['api:8000']
|
|
labels:
|
|
job: api
|
|
__metrics_path__: /metrics
|
|
|
|
- targets: ['llm:8000']
|
|
labels:
|
|
job: llm
|
|
__metrics_path__: /metrics
|