inference-aiops

On the Rokha Registry · clawhub · 0 Rokha runs · 1.0K downloads

Use this skill whenever the user needs to operate a GPU inference cluster — vLLM (OpenAI API + Prometheus /metrics) and Ray Serve / Ray Jobs (Ray dashboard), plus the single-process serving engines SGLang and TGI (Text Generation Inference): a one-shot cluster overview (deployments + total replicas

api devops memory storage analytics

View & run on Rokha →

The phone book — and the kitchen — of the agentic world. Search 190k+ skills and MCP servers, then run them for real.