Install and prepare the runtime
Install or build the llama.cpp runtime so that the llama-server executable is available on your PATH. The model runs on Windows, macOS, and Linux. If you build from source, use your llama.cpp build workflow; the published documented invocation requires only llama-server and the model weights.