llm Local LLM Inference Architecture — From Silicon to Stack
A unified stack for running, routing, and observing LLMs — currently deployed on a laptop with RTX 3050 Ti, with planned expansions to K8s GPU nodes and Apple Silicon clusters.
7 posts
llm A unified stack for running, routing, and observing LLMs — currently deployed on a laptop with RTX 3050 Ti, with planned expansions to K8s GPU nodes and Apple Silicon clusters.
mlx A practical walkthrough of Apple's MLX framework for local LLMs — inference, quantization, fine-tuning on a single Mac, then scaling to multi-Mac clusters with JACCL and Thunderbolt 5.
agents Continuing the kri fleet assistant story: the bugs that shipped after RAG landed, the agentic design we red-teamed before building, a 6-persona audit that found the control plane was broken, and the safety model for letting an LLM act on real infrastructure.
mcp Strip away the hype and an AI agent is one simple thing: a language model running in a loop, calling tools, and reading the results until a goal is reached. A chatbot answers once and stops. An agent acts — it decides its own next step based on what just happened in the environment.
rag A hands-on deep dive into building a Retrieval-Augmented Generation pipeline for a Mac Mini build fleet: data taxonomy, chunking, hybrid retrieval, live-state integration, pgvector, intent routing, grounding, and evaluation.
graphify Claude built a knowledge graph of the kri codebase, documented the token-saving workflow, then immediately ignored it. Here's what happened, what the token data showed, and how a PreToolUse hook now enforces the rule at the machine level.
llm Practical tips to reduce token exhaustion and get more useful responses from AI coding tools.