Gemma-4-12B running at 711 tok/s on AMD GPU

    作者 Gifamoss: AMD

    An AMD engineer just demonstrated something remarkable on camera. The clip captures it all in a single frame: llama.cpp slot logs on the left, nvitop showing a steady 79% GPU utilization, and a localhost chat answering a 26,000-token prompt at 711 tok/s prompt eval for the gemma-4-12B model. You can see the full setup at

    文字记录 (en)

    I