Live connection (MCP) · glama · 0 Rokha runs · 0 downloads · ⚖ MIT License
Eyes for text-only LLMs: decodes screenshots into exact structured text (words, coordinates, sizes, colors) using pure-code CV and OCR. Enables text-only models to reason about UI layouts without vision models or VRAM usage.