llm Local LLM Inference Architecture — From Silicon to Stack
A unified stack for running, routing, and observing LLMs — currently deployed on a laptop with RTX 3050 Ti, with planned expansions to K8s GPU nodes and Apple Silicon clusters.
1 post
llm A unified stack for running, routing, and observing LLMs — currently deployed on a laptop with RTX 3050 Ti, with planned expansions to K8s GPU nodes and Apple Silicon clusters.