A FastAPI service under a full observability stack — metrics · logs · traces · profiles · browser — in one docker compose up, or in Kubernetes.
Quick start · Architecture · Dashboard · Signal to signal · Servers · Logging · Metrics · Configuration · Kubernetes · Development
The handlers have no business logic: they only touch what shows up in Grafana — an in-memory cache, SQLite, an external API (JSONPlaceholder) and the CPU. One application, four servers: Gunicorn · Uvicorn · Hypercorn · Granian.
Needs Docker and just; uv only to work on the code.
just dc up # docker compose up --build -d; just k3d up for Kubernetes
just traffic # RATE=10 DURATION=60 just trafficBoth ways to run it live in deploy/:
compose/docker-compose.yaml and the Kubernetes manifests in
kubernetes/. The just recipes point Compose at its file; by hand it is
docker compose -f deploy/compose/docker-compose.yaml …. The same stack runs in a local Kubernetes
cluster with just k3d up — see Kubernetes.
| what | where |
|---|---|
🐍 the page, the API, /docs |
http://localhost:8000 |
| 📊 Grafana, no login | http://localhost:3000 |
| 🔀 Alloy: components, live debugging | http://localhost:12345 |
Tip
Every button on the page is a trace that starts in the browser. just dc down -v wipes the data;
just run granian runs the application outside Docker against the running stack.
flowchart LR
subgraph sources["Sources"]
direction TB
browser["🌐 Browser<br/>Faro SDK"]
api["🐍 api<br/>a server and its workers"]
docker[("🐳 Docker · ☸️ Kubernetes<br/>container stdout")]
end
alloy{{"Alloy<br/>the only collector"}}
subgraph storage["Storage"]
direction TB
vm[("VictoriaMetrics<br/>metrics")]
loki[("Loki<br/>logs")]
tempo[("Tempo<br/>traces")]
pyroscope[("Pyroscope<br/>profiles")]
end
grafana["📊 Grafana"]
api -- "OTLP: traces, metrics<br/>CPU profiles" --> alloy
api -- "JSON lines" --> docker
docker -- "logs" --> alloy
browser -- "Faro: logs, errors,<br/>Web Vitals, spans" --> alloy
alloy --> vm & loki & tempo & pyroscope
tempo -. "span metrics" .-> vm
storage --> grafana
classDef metrics fill:#B877D9,stroke:#8F3BB8,color:#111
classDef logs fill:#73BF69,stroke:#56A64B,color:#111
classDef traces fill:#5794F2,stroke:#3274D9,color:#111
classDef profiles fill:#FF9830,stroke:#FA6400,color:#111
classDef collector fill:#F55F3E,stroke:#C4162A,color:#fff
classDef ui fill:#F46800,stroke:#C34F00,color:#fff
classDef source fill:#E8E8E8,stroke:#9E9E9E,color:#111
class vm metrics
class loki logs
class tempo traces
class pyroscope profiles
class alloy collector
class grafana ui
class browser,api,docker source
| signal | from the process | through Alloy | stored in |
|---|---|---|---|
| 📈 metrics | OTLP gRPC every 15 s, each worker its own service.instance.id |
otelcol.receiver.otlp → Prometheus model |
VictoriaMetrics, 14 days |
| 📜 logs | one JSON object per line on stdout; the browser's through Faro | loki.source.docker or loki.source.kubernetes · faro.receiver: lvl a label, ids structured metadata |
Loki |
| 🧵 traces | OTLP gRPC; the browser's through Faro | otelcol.receiver.otlp · faro.receiver |
Tempo, 3 days |
| 🔥 profiles | Pyroscope SDK, samples tagged with the root span | pyroscope.receive_http |
Pyroscope |
Note
A memory limiter is first in Alloy's line: under pressure it refuses data instead of the collector being killed. Tempo derives span metrics and Alloy the service graph, both written to VictoriaMetrics.
Alloy draws its own pipeline at http://localhost:12345, with the live rate on every edge: what comes in, which stage it passes, where it goes.
FastAPI: traffic, latency, errors, traces, logs, profile — nothing hardcoded: data sources, service, routes, method, host and thresholds are variables. A series on the route, status code and exception panels links to its traces or log lines — click it; a route in the Routes table narrows the whole dashboard to itself, as Grafana's own operations tables do.
| where | variables |
|---|---|
| the top of the dashboard | Metrics · Logs · Traces · Profiles · Service · Routes matching · Route · Method · Host · Apdex satisfied · Apdex tolerated · Slow trace |
| the Status codes row | Status codes — the classes its panel draws, 4xx and 5xx at first |
| the Latency row | Percentiles — P50, P95 and P99, P95 at first |
| the Workers row | Percentiles — the same, for the workers' latency |
| the Profiling row | Profile type |
| the Traces row | Show — slow or failed at first, slow, failed or all |
| the Logs row | Levels — DEBUG, INFO, WARNING, ERROR, CRITICAL, all at first |
A row's picker reshapes its panel and nothing else: every status class at once, every percentile of the service and of each route.
The order is the order of an investigation: is the service fine, which requests suffer, is the promise kept, why, and the evidence. Groups nest — a dashboard in the v2 schema, the one with rows inside rows.
| group | row | what |
|---|---|---|
| top | RPS · 5xx ratio · P95 now, with sparklines · Apdex · instances · workers · every route with its trend, requests, 2xx, 3xx, 4xx, 5xx, mean, P50, P95 and P99 | |
| Requests | Traffic | RPS by route, the total over them and the total yesterday |
| Status codes | RPS of the classes picked, a total per class and a line per route and code | |
| Latency | the percentiles picked, of the service, yesterday and of each route · the heatmap | |
| Exceptions | by route and type · each message with its type, route and count — from the log | |
| Payload, collapsed | bytes per second · body size P95 by route, requests dashed and responses solid | |
| SLO | error budget left · burn rate over 1 h and 6 h · requests fast enough — all over 7 days, with sparklines · availability against the objective | |
| Runtime | Workers | request share against an even split · the percentiles picked, under their total · in flight, stacked · CPU against one core · involuntary context switches · worker starts |
| Process, collapsed | memory · threads · open files per worker · GC | |
| Profiling, collapsed | flame graph | |
| Traces & logs | Traces | the traces picked in Show, newest first |
| Logs | of the levels picked: lines by level · the stream — time, level, status, duration, request and message in columns |
Every route and every worker is drawn; the variables narrow them — Route to some routes, Host
to the workers of one container or pod. Each title says what is measured, then how it is cut:
RPS — total · by route, CPU — by worker, one core is 100%.
A click on /api/report/{post_id} in Routes, and the dashboard is about that route alone:
In three languages. English, Русский and 中文 are three dashboards in
dashboards/, each translated whole — title, variables, rows,
panels; pick one in the dashboard list or under All dashboards. Grafana translates its own
interface, which follows the browser's language, but never what a dashboard says: a change to one
is a change to all three.
On grafana.com. Its upload takes the classic dashboard JSON, not the v2 schema these are in:
grafana-com/fastapi-observatory.json
is the English one as Grafana itself converts it — GET /apis/dashboard.grafana.app/v1beta1/…/dashboards/observatory-api.
The classic schema has no rows inside rows and no row variables, so the rows lie flat and their
pickers join the variables on top; the data sources stay variables, picked on import.
Lines read the same on every panel. A total is thick, over a light fill, drawn above the rest and
named total — white for the service, or the colour of its class: 2xx green, 3xx blue, 4xx orange,
5xx and exceptions red; P50 blue, P95 white, P99 purple. A thin line without fill is one route or
one worker. A dash is kept for two things: grey for yesterday, white for a reference — the
objective, an even split. Latency zones above 500 ms and 1 s are shaded; under the objective,
availability sits in a red zone. A panel of a few lines of their own — bytes, availability, GC —
draws them at middle width; a table legend lists every series that a panel breaks down.
Thresholds: 5xx over 1% orange, over 5% red; P95 over 500 ms orange, over 1 s red. The availability
target of the budget and the burn rate is the hidden variable slo, 0.995. Burn rate turns red at
3.36× over 1 h and 1.4× over 6 h — the pace that spends 2% and 5% of a 7-day budget in that window.
just traffic fails about 5% of requests on purpose and sends 4% the API refuses (422, 405), so the budget runs out.
A worker is a process: with the GIL it gets about one core, so the CPU panel is in fractions of one core, not of the machine. One worker slower than the others, busier on CPU and preempted more often points at a request that hogs it, not at the service as a whole. Worker starts shows when workers came and went: a deploy, a crash, a recycle, or one Gunicorn killed for missing its 30-second heartbeat.
flowchart LR
L["📜 Logs<br/>Loki"]
T["🧵 Traces<br/>Tempo"]
P["🔥 Profiles<br/>Pyroscope"]
M["📈 Metrics<br/>VictoriaMetrics"]
L -- "trace_id · session_id" --> T
T -- "Logs for this span" --> L
T -- "Profiles for this span" --> P
T -- "span metrics: rate · P95" --> M
M -. "panel links" .-> T
M -. "panel links" .-> L
classDef metrics fill:#B877D9,stroke:#8F3BB8,color:#111
classDef logs fill:#73BF69,stroke:#56A64B,color:#111
classDef traces fill:#5794F2,stroke:#3274D9,color:#111
classDef profiles fill:#FF9830,stroke:#FA6400,color:#111
classDef collector fill:#F55F3E,stroke:#C4162A,color:#fff
classDef ui fill:#F46800,stroke:#C34F00,color:#fff
classDef source fill:#E8E8E8,stroke:#9E9E9E,color:#111
class M metrics
class L logs
class T traces
class P profiles
| from | to | how |
|---|---|---|
| a log line, the browser's included | its trace · every line of its request | trace_id · request_id in structured metadata |
| a browser line | every trace of that browser session | session_id → {span.session.id="…"} |
| a request key from a header or a complaint | its trace | {span.http.response.header.x_request_id="…"} |
| a route in the Routes table | the whole dashboard for that route · its traces | Route set on the same dashboard · TraceQL with the route |
| a route on the traffic or latency panel | its traces · its slow traces | panel link → TraceQL with the route and the Slow trace threshold |
| a status class or a route with its code | the traces with that class · with that code on that route | panel link → TraceQL with the status code |
| an exception on a route | its log lines and stacks · its traces | panel link → LogQL with error_type and route · TraceQL with event.exception.type |
| an exception message | its log lines | panel link → LogQL with the message |
| a span | its logs · its CPU profile · the rate and P95 of its operation | Tempo data source links |
A series carries its links: a click on a bar of /api/fail offers the lines and the traces of
that exception on that route.
① A log line links its trace and every line of its request.
② The line and its trace, side by side — this one started with a click in the page.
③ A span links its logs, its profile and the metrics of its operation.
④ Its profile is the CPU of exactly that request — here one of /api/cpu.
Note
Span profiles label samples by thread. A request that holds the event loop, like /api/cpu,
gets a profile of its own; requests interleaving on the loop may share one.
⑤ A route's line on the latency panel opens its slow traces. ⑥ The service graph is drawn from client and server spans.
⑦ The rate and P95 of a span's operation open beside the trace, from span metrics.
⑧ An exception's line holds its type, message and the whole stack. ⑨ A browser error arrives through Faro with the session it happened in, and the session links every trace of it.
A click on Report on post 3 is one trace — the browser, the API, the cache, JSONPlaceholder, SQLite and the CPU work:
sequenceDiagram
participant B as Browser (Faro)
participant M as AccessMiddleware
participant H as build_report
participant C as MemoryCache
participant J as JSONPlaceholder
participant S as SQLite
B->>M: GET /api/report/3 + traceparent
M->>H: handle, in the server span
par the post
H->>C: cache get (miss)
H->>J: GET /posts/3
H->>S: INSERT
H->>C: cache set
and its comments
H->>J: GET /posts/3/comments
end
H->>H: count primes (profiled)
H-->>M: 200
M->>M: log line with trace_id
M-->>B: 200 + x-request-id
| handler | the trace shows |
|---|---|
GET /api/posts/{id} |
cache → on a miss JSONPlaceholder → INSERT → cache |
GET /api/posts |
a SELECT |
POST /api/posts |
an INSERT — the only request with a body |
GET /api/cpu?below=N |
count primes and its profile |
GET /api/report/{id} |
all of the above, the fetches in parallel |
GET /api/posts/1000 |
a 404 from the source |
GET /api/fail?kind=… |
a 500 of the kind asked for — runtime, invalid, lookup, timeout, permission, each its own exception type: the exception on the span, the stack in the log |
GET /api/posts/latest · GET /api/cpu?below= past the limit |
a 422: FastAPI refuses the parameter before the handler runs |
DELETE /api/posts/{id} |
a 405: the route exists, the method does not |
OBSERVATORY__SERVER__KIND picks the server; every one implements servers.Server and loads the
same factory, app:create_app, in each process that serves.
| server | workers | process model | a worker killed with SIGKILL |
|---|---|---|---|
| Gunicorn + uvicorn-worker | forked by the master | inherits memory, not threads | replaced |
| Uvicorn | spawned by a supervisor | a new interpreter | replaced |
| Hypercorn | spawned by a master | a new interpreter | |
| Granian | spawned by a master | a new interpreter | replaced |
sequenceDiagram
autonumber
participant M as master / supervisor
participant W as each serving process
M->>M: Settings() · configure_logging()
M->>M: prepare_database() — once
M->>W: fork or spawn
Note over W: forked: memory without threads<br/>spawned: nothing at all
W->>W: create_app(): logging · tracing · metrics · profiling
W->>W: open SQLite, serve
Important
create_app() assumes nothing set up before it, so the same code works forked and spawned.
Telemetry starts threads, and a thread does not survive a fork — that is why Gunicorn's
preload_app stays off: with it, profiles are lost and Python warns about forking a
multi-threaded process.
Checked — each server with two workers, the same traffic:
| Gunicorn | Uvicorn | Hypercorn | Granian | |
|---|---|---|---|---|
| requests served | ✅ | ✅ | ✅ | ✅ |
| every log line JSON | ✅ | ✅ | ✅ | ✅ |
access lines with trace_id |
✅ | ✅ | ✅ | ✅ |
| workers sending telemetry | 2 | 2 | 2 | 2 |
| process metrics | ✅ | ✅ | ✅ | ✅ |
| traces and profiles | ✅ | ✅ | ✅ | ✅ |
flowchart LR
record["logger.info('post created',<br/>extra={'post_id': 8})"]
factory["record factory<br/>+ request_id · trace_id · span_id"]
fmt["JSON formatter<br/>on the root's handler"]
out["stdout"]
alloy["Alloy<br/>lvl → label<br/>ids → structured metadata"]
loki[("Loki")]
record --> factory --> fmt --> out --> alloy --> loki
classDef metrics fill:#B877D9,stroke:#8F3BB8,color:#111
classDef logs fill:#73BF69,stroke:#56A64B,color:#111
classDef traces fill:#5794F2,stroke:#3274D9,color:#111
classDef profiles fill:#FF9830,stroke:#FA6400,color:#111
classDef collector fill:#F55F3E,stroke:#C4162A,color:#fff
classDef ui fill:#F46800,stroke:#C34F00,color:#fff
classDef source fill:#E8E8E8,stroke:#9E9E9E,color:#111
class loki logs
class alloy collector
class record,factory,fmt,out source
{"ts":"2026-10-03T10:19:44.631+00:00","lvl":"INFO","msg":"HTTP request handled","logger":"observatory","caller":"middleware:_write:70","request_id":"7cefc9e7…","trace_id":"ef23e199…","span_id":"…","method":"GET","path":"/api/posts/3","route":"/api/posts/{post_id}","status":200,"duration_ms":297}| who writes | logger | level |
|---|---|---|
| the application, the telemetry setup | observatory |
OBSERVATORY__LOG_LEVEL |
| Gunicorn · Uvicorn · Hypercorn · Granian | their own, formatted as JSON | OBSERVATORY__LOG_LEVEL |
| libraries: httpx, OpenTelemetry, … | their own | the root's WARNING |
The access line's level follows the status: INFO, WARNING for a 4xx, ERROR for a 5xx. An
unhandled exception becomes error_type, error_message, error_stack. route is the matched template,
the one the metrics carry as http_route, so a line and its metrics group alike. AccessMiddleware
writes one line per response inside the request span and returns x-request-id, which the FastAPI
instrumentation also records on the span; the servers' access logs are off.
Browser lines arrive through Faro. Every one carries its session_id, and those of the page's own
HTTP calls their trace_id and span_id — all three in structured metadata, like the
application's keys.
How every line becomes JSON
configure_logging(level) runs in every process that writes lines — main() and create_app():
it gives the root a stdout handler, puts the JSON formatter on it and sets observatory's level.
A repeat call stacks nothing.
- Levels are checked on the logger called, not on its ancestors: the root stays at
WARNING, yetobservatory'sINFOreaches the root's handler. observatorykeepspropagate = True— withFalseits lines would never reach that handler.- Servers with handlers of their own get them formatted too (
GunicornJsonLogger); the others are pointed at the root.
The process ships only the server's and its own metrics: a view in configure_metrics keeps
http.server.*, process.* and cpython.*. httpx and SQLite still give spans, not metrics.
Note
VictoriaMetrics turns OpenTelemetry names into Prometheus ones: dots become _, the unit a
suffix — http.server.request.duration (s) → http_server_request_duration_seconds.
Alloy's otelcol.exporter.prometheus turns OTLP into Prometheus series by the
OpenTelemetry → Prometheus mapping;
each of its options is written out in observability/alloy/receiver.alloy.
| on the series | from |
|---|---|
| the name | the metric name, . → _, with the unit and _total for counters |
job |
resource service.name |
instance |
resource service.instance.id — <host>-<pid>, one per worker |
| every other label | the data point's attributes, . → _ |
target_info{job, instance, …} |
one series per worker with the rest of the resource: host_name, deployment_environment_name, the SDK — joined on job and instance when needed |
otel_scope_name and otel_scope_version are switched on but do not arrive yet: Alloy 1.20.1
ignores include_scope_labels (grafana/alloy#6787).
| metric | from | its own labels |
|---|---|---|
http_server_request_duration_seconds |
FastAPI | http_route, http_request_method, http_response_status_code, error_type, url_scheme, network_protocol_version |
http_server_active_requests |
FastAPI | http_request_method, url_scheme |
http_server_request_body_size_bytes, http_server_response_body_size_bytes |
FastAPI | as the duration |
process_cpu_time_seconds_total, process_memory_usage_bytes, process_thread_count, process_open_file_descriptor_count, … |
each worker about itself | — |
cpython_gc_collections_total, cpython_gc_collected_objects_total, … |
each worker's GC | generation and cpython_gc_generation — the instrumentation sends both |
traces_spanmetrics_calls_total, traces_spanmetrics_latency |
Tempo, from spans of both services | service, span_name, span_kind, status_code, the dimensions of tempo.yaml, source="tempo", Tempo's __metrics_gen_instance |
traces_service_graph_request_* |
Alloy, from span pairs | client, server, connection_type, failed, virtual_node |
Sources: OpenTelemetry semantic conventions · FastAPI instrumentation · system metrics · span metrics.
Environment variables with the OBSERVATORY__ prefix; groups nest with __.
| variable | default |
|---|---|
OBSERVATORY__SERVER__KIND |
required: gunicorn · uvicorn · hypercorn · granian |
OBSERVATORY__SERVER__HOST · PORT · WORKERS |
127.0.0.1 · 8000 · 2 |
OBSERVATORY__LOG_LEVEL |
info |
OBSERVATORY__DB__PATH |
observatory.db |
OBSERVATORY__OBS__SERVICE_NAME · ENVIRONMENT |
api · development — or production, staging, test |
OBSERVATORY__OBS__OTLP__ENDPOINT |
http://localhost:4317 |
OBSERVATORY__OBS__PYROSCOPE__URL |
http://localhost:4040 |
OBSERVATORY__OBS__FARO__COLLECTOR_URL |
http://localhost:12347/collect — as the browser sees it |
Each server's own options
Under the same OBSERVATORY__SERVER__ prefix; another server's option is refused at startup.
KIND |
options |
|---|---|
gunicorn |
TIMEOUT · GRACEFUL_TIMEOUT · KEEPALIVE · MAX_REQUESTS · MAX_REQUESTS_JITTER · BACKLOG |
uvicorn |
TIMEOUT_KEEP_ALIVE · TIMEOUT_GRACEFUL_SHUTDOWN · LIMIT_CONCURRENCY · LIMIT_MAX_REQUESTS · BACKLOG |
hypercorn |
KEEP_ALIVE_TIMEOUT · GRACEFUL_TIMEOUT · MAX_REQUESTS · MAX_REQUESTS_JITTER · BACKLOG |
granian |
RUNTIME_THREADS · BLOCKING_THREADS · BACKPRESSURE · BACKLOG · WORKERS_LIFETIME · WORKERS_KILL_TIMEOUT |
MAX_REQUESTS=1000 MAX_REQUESTS_JITTER=100 under Gunicorn restarts each worker after 1000–1100
requests.
just k3d up runs the same stack in a local k3d cluster: it builds the image,
imports it, applies deploy/kubernetes/ and waits for the rollout; just k3d down deletes the cluster. The ports are Compose's — 3000, 8000, 12345, 12347 — so stop one
before starting the other.
Nothing is copied. Kustomize builds the ConfigMaps from the files Compose mounts, and the Services
carry the Compose service names, so loki:3100 and alloy:4317 mean the same on both. The
dashboard does not know the platform either: it asks only for service.name and
service.instance.id.
| Compose | Kubernetes | |
|---|---|---|
| Alloy reads | receiver.alloy + docker.alloy |
receiver.alloy + kubernetes.alloy |
| container output | loki.source.docker, through docker.sock |
loki.source.kubernetes, through the API; a Role reads the pods and their logs in one namespace |
a log line's service_name |
the Compose service | the pod's app.kubernetes.io/name |
an instance — host_name |
the container | the pod; the API runs two, of two workers each |
| where a process runs | — | k8s.pod.name, k8s.pod.uid, k8s.namespace.name, k8s.node.name from the downward API, in OTEL_RESOURCE_ATTRIBUTES |
Note
Alloy reads pod output through the API server, which is enough for one small cluster. A large one
reads the nodes' log files from a DaemonSet, as Grafana's k8s-monitoring chart does, and runs
receiver.alloy in a Deployment of its own.
just is the command line: a command, then what it acts on. A value it does not know is an error
that lists the ones it does; just alone lists the commands.
| command | does |
|---|---|
just install · just hooks |
the venv · git hooks via prek |
just lint |
every check — what CI would run |
just lint ruff · flake8 · mypy |
ruff format and check · wemake-python-styleguide · mypy |
just lint slotscheck · spelling · pyproject |
slotscheck · codespell and typos · pyproject-fmt |
just dc up · down · ps · logs |
the stack in Compose; just docker-compose … is the same |
just k3d up · down · ps · logs |
the stack in a k3d cluster |
just dc logs api grafana · just k3d logs api |
the logs of chosen services only |
just traffic |
requests at the running stack, either one: RATE, DURATION, and the mix — CPU_PERCENT, REPORT_PERCENT, CPU_BELOW_MAX, FAIL_PERCENT, INVALID_PERCENT; API_WORKERS=8 just dc up for a heavier one |
just run gunicorn |
the application outside Docker, against the running stack |
| version | port | role | |
|---|---|---|---|
| Grafana | 13.2.3 | 3000 |
dashboards, Explore, Drilldown |
| Alloy | 1.20.1 | 12345 · 12347 |
the only collector · the Faro receiver |
| VictoriaMetrics | 1.153.0 | 8428 |
metrics |
| Loki | 3.7.8 | 3100 |
logs |
| Tempo | 3.1.0 | 3200 |
traces, span metrics |
| Pyroscope | 2.3.1 | 4040 |
CPU profiles |
| Faro Web SDK | 2.12.1 | — | the browser's telemetry |
| Gunicorn · Uvicorn · Hypercorn · Granian | 26.2 · 0.54 · 0.18 · 2.8.4 | 8000 |
the servers; Gunicorn in Docker |
Everything is published on 127.0.0.1 only.


























