Files
stack/infra/images/api.Dockerfile
kert 1843c0a9a0
All checks were successful
CI / lint (push) Successful in 33s
CI / test (push) Successful in 2m6s
Deploy / notebooks (push) Has been skipped
CI / notebooks-smoke (push) Successful in 1m31s
Deploy / zotero (push) Has been skipped
Deploy / docs (push) Has been skipped
Deploy / llm (push) Successful in 1m21s
Deploy / mc (push) Has been skipped
Deploy / api (push) Successful in 1m46s
Infra CI / docs (push) Successful in 21s
Infra CI / llm (push) Successful in 15s
Infra CI / mc (push) Successful in 14s
Deploy / report (push) Successful in 11s
Infra CI / zotero (push) Successful in 17s
Infra CI / notebooks (push) Successful in 50s
Infra CI / api (push) Successful in 19s
fix(perf,infra): actually export the metrics — /metrics route, perf extra in the api/llm images, scrape targets, telemetry on (refs #579)
The P26 telemetry path never reached Prometheus: setup_meter_provider
only installs an in-process PrometheusMetricReader, nothing served the
registry (api and llm answered 404 on /metrics), the images never
installed the perf extra, STACK_TELEMETRY was off, and no scrape target
existed — so the data-pipelines request-rate panel was empty from the
day it was written. Now: perf.middleware.instrument mounts GET /metrics
(prometheus_client registry), the api and llm images install --extra
perf, compose sets STACK_TELEMETRY=true on both, services.yml scrapes
api:8000 and llm:8000, and both dashboards' request-rate panels query
http_server_duration_milliseconds_count (what the FastAPI instrumentor
emits; stack_http_server_requests_total never existed). Verified live:
stack_llm_dispatch_total is queryable in Prometheus with job=llm.
2026-09-11 18:59:40 -04:00

45 lines
1.7 KiB
Docker

# syntax=docker/dockerfile:1
FROM ghcr.io/astral-sh/uv:python3.13-bookworm-slim
WORKDIR /app
# Local package registry (Gitea) — set via --build-arg to pull from mirror
ARG PYPI_INDEX_URL=""
# Patch base image CVEs + install curl for healthcheck.
# (Python cold-start on this image is 10-13s — too slow for the 5s
# healthcheck timeout, so Docker marks the container unhealthy even
# though /health responds in milliseconds.)
RUN apt-get update && apt-get upgrade -y && apt-get install -y --no-install-recommends curl && rm -rf /var/lib/apt/lists/*
# Copy project files for install
COPY pyproject.toml uv.lock README.md ./
COPY src/ src/
# Install the package (no dev deps). The server needs the cli extra
# bundle (aco + api + bib + mail): uvicorn/fastapi/pyjwt live in the
# api extra, and /health imports aco.pipe + bib.store + duckdb — a
# bare `uv sync --no-dev` installs none of them (base deps only),
# which shipped an image that crash-looped on `Failed to spawn:
# uvicorn`.
ENV UV_PYTHON_PREFERENCE=only-system \
UV_LINK_MODE=copy \
UV_PROJECT_ENVIRONMENT=.venv \
UV_INDEX_URL=${PYPI_INDEX_URL}
RUN uv sync --no-dev --extra cli --extra perf && uv pip install -e .
# Config
COPY stack.toml ./
EXPOSE 8000
# --max-time bounds curl itself: docker's timeout only stops *waiting* —
# the probe process lives on, and a wedged server once accumulated ~1500
# hung healthcheck zombies this way.
HEALTHCHECK --interval=30s --timeout=5s --retries=3 \
CMD curl -sf --max-time 4 http://localhost:8000/health || exit 1
CMD ["uv", "run", "--no-sync", "uvicorn", "api.server:app", \
"--host", "0.0.0.0", "--port", "8000", \
"--workers", "1", "--log-level", "info"]