krioperationsdevops

Operating kri — The Fleet Manager's Field Guide

Operating kri — The Fleet Manager's Field Guide

DevOps Field Guide


You’ve deployed kri. Postgres and Redis are in Docker, the FastAPI backend is on port 8000, and the React dashboard is on 5173. This post is the practical reference for the first week of operating it: starting services, reading logs, and diagnosing a node that won’t bootstrap.

1. Starting kri

Everything goes through ./scripts/kri.sh. It handles startup order, PID files, health checks on the Docker containers, and log routing.

Command reference

CommandWhat it does
./scripts/kri.sh startBrings up postgres + redis (Docker), waits for healthy, then starts API, Celery worker, and Vite frontend in that order
./scripts/kri.sh stopKills frontend, worker, API (reverse order), then calls docker compose stop
./scripts/kri.sh restartstop + start
./scripts/kri.sh statusReads each PID file, checks if the process is alive, and prints Docker container health
./scripts/kri.sh logs api|worker|frontendtail -f on the requested log file — Ctrl-C to exit
./scripts/kri.sh test [grep]Runs the Playwright E2E suite; pass an optional grep pattern to filter specs
# Normal startup sequence
./scripts/kri.sh start

# Expected output
  kri fleet management platform
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
✓ Infrastructure ready
✓ API running on :8000 (log: .kri-logs/api.log)
✓ Celery worker running (log: .kri-logs/worker.log)
✓ Frontend running on :5173 (log: .kri-logs/frontend.log)

✓ kri is up →  http://localhost:5173
   API     →  http://localhost:8000/docs

ℹ

The script waits up to 20 seconds for both Docker containers to report healthy before starting the application services. If they stay unhealthy, run docker compose -f deploy/docker-compose.yml ps to investigate.

Under the hood — PID tracking

Each process writes its PID to .kri-pids/<name>.pid. The status command reads these files and runs kill -0 <pid> to confirm the process is alive. If kri.sh crashes mid-start, stale PID files can cause the next start to emit false “already running” warnings — just delete .kri-pids/ and try again.

2. Where Logs Live

ServiceLog pathWhat to look for
API.kri-logs/api.logHTTP requests, SQL queries, auth events, FastAPI startup errors
Worker.kri-logs/worker.logBootstrap task lifecycle, Ansible stdout lines, drift computation, SBOM tasks
Frontend.kri-logs/frontend.logVite dev server output, HMR reload events, TypeScript compile errors

Useful log commands

# Use the kri.sh shortcut
./scripts/kri.sh logs worker

# Or tail directly
tail -f .kri-logs/api.log
tail -f .kri-logs/worker.log

# Filter bootstrap-related lines
grep -i "bootstrap" .kri-logs/worker.log | tail -50

# Surface errors in the last 200 lines of the API log
tail -200 .kri-logs/api.log | grep -i "error\|exception\|traceback"

# Live filter — only worker lines mentioning a task
tail -f .kri-logs/worker.log | grep --line-buffered "Task\|TASK\|bootstrap"

Log and telemetry architecture

Application Layer Persistent Store File Outputs FastAPI :8000 HTTP / auth / routes Celery Worker bootstrap_node task Vite Frontend :5173 React / HMR dev server PostgreSQL nodes.bootstrap_status nodes.bootstrap_logs (stdout) audit_events every user action logged Redis Celery broker + results .kri-logs/api.log .kri-logs/worker.log .kri-logs/frontend.log /srv/salt/pillar/ <minion_id>.sls

3. Debugging a Stuck Bootstrap

A bootstrap runs as a Celery task that SSHs into the target Mac Mini via Ansible, installs the Salt minion, writes a pillar file, and waits for the minion to check in. Failures at any stage surface through the API.

The bootstrap flow

FastAPI POST /bootstrap enqueues task Celery Worker bootstrap_node() reads Settings ansible-runner SSH → Mac Mini install salt pkg Mac Mini salt-minion starts connects to salt-master ① queued ② task starts ③ SSH + install ④ minion online status written to DB at each step

Common failures at a glance

SymptomLikely causeFix
UNREACHABLEWrong IP or Mac Mini is offVerify IP; test with ping
Authentication failureWrong SSH credentialsSettings → Bootstrap — update SSH username / password
Timeout at 20 minSalt minion installed but cannot reach masterCheck network path; verify Salt master address in Settings
no hosts matchedminion_id contains illegal charactersOnly [a-zA-Z0-9._-] allowed, max 128 chars
ansible.posix not foundMissing Ansible collectionansible-galaxy collection install ansible.posix
Stuck in bootstrappingCelery worker crashed mid-taskCancel via API then retry (see below)

Step-by-step debug procedure

  1. Find the node in the database. The minion_id is the Salt identifier — usually hostname.local or hostname.domain.

    docker exec deploy-postgres-1 psql -U fleet -d fleet_demo -c \
      "SELECT id, minion_id, bootstrap_status, bootstrap_error, bootstrap_ip \
       FROM nodes WHERE minion_id LIKE '%mm1%';"
  2. Check the worker log for the last Ansible task name — the worker writes [blocked at: TASK <name>] into bootstrap_error every 5 seconds while running.

    grep -i "bootstrap\|blocked\|TASK" .kri-logs/worker.log | tail -30
  3. Pull the full bootstrap log from the API. This returns the complete Ansible stdout stored in the database after the run ends.

    curl -s http://localhost:8000/api/v1/ansible/bootstrap/<node_id>/logs \
      -H "Authorization: Bearer <token>" \
      | jq '{status: .bootstrap_status, stdout: .ansible_stdout}'
  4. Cancel a stuck bootstrap if the status is bootstrapping or pending and the worker log has gone silent. This resets the status to failed so you can retry.

    curl -s -X POST http://localhost:8000/api/v1/ansible/bootstrap/<node_id>/cancel \
      -H "Authorization: Bearer <token>"
  5. Fix the root cause, then trigger a fresh bootstrap from Fleet → Add Node. Re-submitting with the same minion_id is safe — kri reuses the existing node row.

4. Reading Bootstrap Telemetry

kri stores fine-grained telemetry for every bootstrap attempt. You don’t need to log into the Mac Mini to know what happened.

What kri captures

DataDB columnNotes
Current statusnodes.bootstrap_statuspending → bootstrapping → completed or failed
Target IPnodes.bootstrap_ipIP used for the Ansible SSH connection
Error / last tasknodes.bootstrap_errorHuman-readable failure reason, or [blocked at: TASK ] while running
Full Ansible stdoutnodes.bootstrap_logsStored after the run ends (success or failure); exposed by the /logs endpoint
User actionsaudit_events tableEvery bootstrap trigger, cancel, and Settings change is recorded with actor + timestamp

API endpoints

GET /api/v1/ansible/bootstrap/{node_id}/status

Lightweight poll — returns status, bootstrap_ip, and bootstrap_error. No stdout. Use this to track progress in a loop.

GET /api/v1/ansible/bootstrap/{node_id}/logs

Full detail — returns status, complete Ansible stdout (ansible_stdout), pillar file path, and pillar file contents. Use this for post-mortem on failures.

# Extract just the useful fields from the logs endpoint
curl -s http://localhost:8000/api/v1/ansible/bootstrap/<node_id>/logs \
  -H "Authorization: Bearer <token>" \
  | jq '{
      status:     .bootstrap_status,
      last_error: (.ansible_stdout // "no stdout yet"),
      pillar_ok:  (.pillar | startswith("(") | not)
    }'

# Poll status until it leaves 'bootstrapping'
watch -n5 'curl -s http://localhost:8000/api/v1/ansible/bootstrap/<node_id>/status \
  -H "Authorization: Bearer <token>" | jq .bootstrap_status'

Enabling verbose Ansible output

By default Ansible runs at verbosity 0 (minimal output). To get SSH connection details, module arguments, and return values in the worker log, set ANSIBLE_VERBOSITY=2 in your .env file and restart the worker. Use level 4 for full debug output (very noisy).

# .env — add or update this line
ANSIBLE_VERBOSITY=2

# Apply by restarting
./scripts/kri.sh restart

⚠

Verbosity 2 and above will log SSH credentials and host details to .kri-logs/worker.log. Rotate or delete the log after debugging.

Querying audit events directly

docker exec deploy-postgres-1 psql -U fleet -d fleet_demo -c \
  "SELECT action, actor, resource_type, event_at \
   FROM audit_events ORDER BY event_at DESC LIMIT 20;"

Enjoyed this post?

Get the next one in your inbox — only when I ship something worth reading.

Newsletter form not configured.

Or follow on Substack for the newsletter.

Comments via GitHub Discussions

Comments not configured. Set GISCUS env vars to enable.