Operating kri — The Fleet Manager's Field Guide
DevOps Field Guide
You’ve deployed kri. Postgres and Redis are in Docker, the FastAPI backend is on port 8000, and the React dashboard is on 5173. This post is the practical reference for the first week of operating it: starting services, reading logs, and diagnosing a node that won’t bootstrap.
1. Starting kri
Everything goes through ./scripts/kri.sh. It handles startup order, PID files, health checks on the Docker containers, and log routing.
Command reference
| Command | What it does |
|---|---|
| ./scripts/kri.sh start | Brings up postgres + redis (Docker), waits for healthy, then starts API, Celery worker, and Vite frontend in that order |
| ./scripts/kri.sh stop | Kills frontend, worker, API (reverse order), then calls docker compose stop |
| ./scripts/kri.sh restart | stop + start |
| ./scripts/kri.sh status | Reads each PID file, checks if the process is alive, and prints Docker container health |
| ./scripts/kri.sh logs api|worker|frontend | tail -f on the requested log file — Ctrl-C to exit |
| ./scripts/kri.sh test [grep] | Runs the Playwright E2E suite; pass an optional grep pattern to filter specs |
# Normal startup sequence
./scripts/kri.sh start
# Expected output
kri fleet management platform
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
✓ Infrastructure ready
✓ API running on :8000 (log: .kri-logs/api.log)
✓ Celery worker running (log: .kri-logs/worker.log)
✓ Frontend running on :5173 (log: .kri-logs/frontend.log)
✓ kri is up → http://localhost:5173
API → http://localhost:8000/docs
ℹ
The script waits up to 20 seconds for both Docker containers to report healthy before starting the application services. If they stay unhealthy, run docker compose -f deploy/docker-compose.yml ps to investigate.
Under the hood — PID tracking
Each process writes its PID to .kri-pids/<name>.pid. The status command reads these files and runs kill -0 <pid> to confirm the process is alive. If kri.sh crashes mid-start, stale PID files can cause the next start to emit false “already running” warnings — just delete .kri-pids/ and try again.
2. Where Logs Live
| Service | Log path | What to look for |
|---|---|---|
| API | .kri-logs/api.log | HTTP requests, SQL queries, auth events, FastAPI startup errors |
| Worker | .kri-logs/worker.log | Bootstrap task lifecycle, Ansible stdout lines, drift computation, SBOM tasks |
| Frontend | .kri-logs/frontend.log | Vite dev server output, HMR reload events, TypeScript compile errors |
Useful log commands
# Use the kri.sh shortcut
./scripts/kri.sh logs worker
# Or tail directly
tail -f .kri-logs/api.log
tail -f .kri-logs/worker.log
# Filter bootstrap-related lines
grep -i "bootstrap" .kri-logs/worker.log | tail -50
# Surface errors in the last 200 lines of the API log
tail -200 .kri-logs/api.log | grep -i "error\|exception\|traceback"
# Live filter — only worker lines mentioning a task
tail -f .kri-logs/worker.log | grep --line-buffered "Task\|TASK\|bootstrap"
Log and telemetry architecture
3. Debugging a Stuck Bootstrap
A bootstrap runs as a Celery task that SSHs into the target Mac Mini via Ansible, installs the Salt minion, writes a pillar file, and waits for the minion to check in. Failures at any stage surface through the API.
The bootstrap flow
Common failures at a glance
| Symptom | Likely cause | Fix |
|---|---|---|
| UNREACHABLE | Wrong IP or Mac Mini is off | Verify IP; test with ping |
| Authentication failure | Wrong SSH credentials | Settings → Bootstrap — update SSH username / password |
| Timeout at 20 min | Salt minion installed but cannot reach master | Check network path; verify Salt master address in Settings |
| no hosts matched | minion_id contains illegal characters | Only [a-zA-Z0-9._-] allowed, max 128 chars |
| ansible.posix not found | Missing Ansible collection | ansible-galaxy collection install ansible.posix |
| Stuck in bootstrapping | Celery worker crashed mid-task | Cancel via API then retry (see below) |
Step-by-step debug procedure
-
Find the node in the database. The minion_id is the Salt identifier — usually
hostname.localorhostname.domain.docker exec deploy-postgres-1 psql -U fleet -d fleet_demo -c \ "SELECT id, minion_id, bootstrap_status, bootstrap_error, bootstrap_ip \ FROM nodes WHERE minion_id LIKE '%mm1%';" -
Check the worker log for the last Ansible task name — the worker writes
[blocked at: TASK <name>]intobootstrap_errorevery 5 seconds while running.grep -i "bootstrap\|blocked\|TASK" .kri-logs/worker.log | tail -30 -
Pull the full bootstrap log from the API. This returns the complete Ansible stdout stored in the database after the run ends.
curl -s http://localhost:8000/api/v1/ansible/bootstrap/<node_id>/logs \ -H "Authorization: Bearer <token>" \ | jq '{status: .bootstrap_status, stdout: .ansible_stdout}' -
Cancel a stuck bootstrap if the status is
bootstrappingorpendingand the worker log has gone silent. This resets the status tofailedso you can retry.curl -s -X POST http://localhost:8000/api/v1/ansible/bootstrap/<node_id>/cancel \ -H "Authorization: Bearer <token>" -
Fix the root cause, then trigger a fresh bootstrap from Fleet → Add Node. Re-submitting with the same minion_id is safe — kri reuses the existing node row.
4. Reading Bootstrap Telemetry
kri stores fine-grained telemetry for every bootstrap attempt. You don’t need to log into the Mac Mini to know what happened.
What kri captures
| Data | DB column | Notes |
|---|---|---|
| Current status | nodes.bootstrap_status | pending → bootstrapping → completed or failed |
| Target IP | nodes.bootstrap_ip | IP used for the Ansible SSH connection |
| Error / last task | nodes.bootstrap_error | Human-readable failure reason, or [blocked at: TASK |
| Full Ansible stdout | nodes.bootstrap_logs | Stored after the run ends (success or failure); exposed by the /logs endpoint |
| User actions | audit_events table | Every bootstrap trigger, cancel, and Settings change is recorded with actor + timestamp |
API endpoints
GET /api/v1/ansible/bootstrap/{node_id}/status
Lightweight poll — returns status, bootstrap_ip, and bootstrap_error. No stdout. Use this to track progress in a loop.
GET /api/v1/ansible/bootstrap/{node_id}/logs
Full detail — returns status, complete Ansible stdout (ansible_stdout), pillar file path, and pillar file contents. Use this for post-mortem on failures.
# Extract just the useful fields from the logs endpoint
curl -s http://localhost:8000/api/v1/ansible/bootstrap/<node_id>/logs \
-H "Authorization: Bearer <token>" \
| jq '{
status: .bootstrap_status,
last_error: (.ansible_stdout // "no stdout yet"),
pillar_ok: (.pillar | startswith("(") | not)
}'
# Poll status until it leaves 'bootstrapping'
watch -n5 'curl -s http://localhost:8000/api/v1/ansible/bootstrap/<node_id>/status \
-H "Authorization: Bearer <token>" | jq .bootstrap_status'
Enabling verbose Ansible output
By default Ansible runs at verbosity 0 (minimal output). To get SSH connection details, module arguments, and return values in the worker log, set ANSIBLE_VERBOSITY=2 in your .env file and restart the worker. Use level 4 for full debug output (very noisy).
# .env — add or update this line
ANSIBLE_VERBOSITY=2
# Apply by restarting
./scripts/kri.sh restart
⚠
Verbosity 2 and above will log SSH credentials and host details to .kri-logs/worker.log. Rotate or delete the log after debugging.
Querying audit events directly
docker exec deploy-postgres-1 psql -U fleet -d fleet_demo -c \
"SELECT action, actor, resource_type, event_at \
FROM audit_events ORDER BY event_at DESC LIMIT 20;" Enjoyed this post?
Get the next one in your inbox — only when I ship something worth reading.
Newsletter form not configured.
Or follow on Substack for the newsletter.
Comments via GitHub Discussions
Comments not configured. Set GISCUS env vars to enable.