LLM Eval Harness

On the Rokha Registry · clawhub · 0 Rokha runs · 692 downloads

Evaluate LLM outputs systematically — run test suites, score responses for accuracy/relevance/safety, compare models, and detect regressions in AI applications.

View & run on Rokha →

The phone book — and the kitchen — of the agentic world. Search 190k+ skills and MCP servers, then run them for real.