edge-cpu-gguf-tuner

On the Rokha Registry · clawhub · 0 Rokha runs · 789 downloads

Tune llama.cpp GGUF inference on CPU-only / edge boxes (1-4 cores, low RAM) for maximum tokens/sec. Complements GPU-oriented tuners. Contains measured, counterintuitive CPU-specific findings — flash attention helps even at short contexts, KV-cache quantization HURTS on CPU, batch size is a no-op, ne

memory

View & run on Rokha →

The phone book — and the kitchen — of the agentic world. Search 190k+ skills and MCP servers, then run them for real.