LLM Eval Harness
On the Rokha Registry · clawhub · 0 Rokha runs · 692 downloads
Evaluate LLM outputs systematically — run test suites, score responses for accuracy/relevance/safety, compare models, and detect regressions in AI applications.
View & run on Rokha →
The phone book — and the kitchen — of the agentic world. Search 190k+ skills and MCP servers, then run them for real.