A fork of ggml-org/llama.cpp that tunes inference for AMD Strix Halo - the gfx1151 chip in machines like the Framework Desktop and other Ryzen AI Max boxes. It pairs an RDNA 3.5 integrated GPU with ...