New `llm.metrics`: one lazily-built instrument set off `perf.meter("stack.llm")`,
so every call site is a no-op mock when telemetry is disabled (as it is in CI and,
today, in the container — see below). A build lock guards first use because
`embed_texts` fans batches across a thread pool.
Instruments, all `stack_llm_…` per the P26 convention:
stack_llm_dispatch_total{host,kind} pool slots handed out
stack_llm_inflight{host} slots held right now
stack_llm_dispatch_seconds{host,kind} how long a slot was held
stack_llm_embedded_texts_total{host} texts embedded
stack_llm_embed_batch_seconds{host} one /api/embed round trip
stack_llm_chat_seconds{stage} retrieve | generate
stack_llm_chat_total{outcome} ok | error
stack_llm_chat_tokens_total{direction} prompt | completion (Ollama's
final frame; skipped when absent)
stack_llm_indexed_chunks_total{collection}
stack_llm_indexed_items_total{collection,outcome} indexed | skipped
Attribute values are closed sets — host, kind, stage, outcome, collection — never
item keys, dockets or question text, so the series count stays bounded.
Call sites: `HostPool.acquire`/`acquire_generation` wrap the held slot (kind embed
vs generate); `embed_texts` times each batch around the POST. `stream_answer` is
now a thin wrapper that counts the outcome once (ok, or error and re-raise) around
`_stream_answer`, which records the retrieve/generate stage split and the token
counts. The indexer counts one item at each of its five existing decision points;
no behaviour changed — every added line is a metric call.
The #575 tagging chain does not exist yet, so tag throughput and abstain rate from
the issue's scope are deliberately left out; they belong with that module.
Dashboard `infra/grafana/dashboards/llm.json` (uid `llm`, "LLM service",
schemaVersion 39): request rate by status filtered to `service_name="llm"`, chat
outcomes, dispatch rate and in-flight per host, p50/p95 dispatch latency by
host+kind, chat latency by stage, embedding throughput, chat tokens, indexed
chunks and items per collection, plus a Loki tail.
Export path — the metrics do NOT reach Prometheus yet, and this commit does not
change that. `perf._meter.setup_meter_provider` installs only a
`PrometheusMetricReader` (an in-process prometheus_client registry); the
`OTEL_EXPORTER_OTLP_ENDPOINT` the llm container sets is read by `_tracer.py` for
spans only, so nothing is pushed to otel-collector:4317 and nothing is exposed for
a scrape (`llm` has no /metrics route and no target in
infra/prometheus/targets/services.yml). On top of that `[telemetry] enabled` is
false in stack.toml and the llm service sets no STACK_TELEMETRY, so `perf.meter`
hands back the no-op mock today. Making #579's "metrics scraped" true needs a
follow-up that (a) enables telemetry for llm, (b) gives metrics a real exporter,
and (c) accounts for the collector's `namespace: stack`, which would re-prefix
OTLP-delivered names to `stack_stack_llm_…`. Filed as a note on #579.
Tests: tests/llm/test_metrics.py patches `perf.meter` with a recording fake and
asserts every instrument name, kind and attribute set, drives the real pool, chat
and indexer paths through it, and checks the helpers stay inert against the real
`MockMeter`; tests/test_observability.py gains the llm dashboard to
EXPECTED_DASHBOARDS plus panel/datasource/quantile assertions.