Large Model Visual Question Answering Skill | 大模型视觉问答技能
On the Rokha Registry · clawhub · 0 Rokha runs · 1.8K downloads
Conducts open-ended Q&A on image content based on computer vision and large language models, supporting any questions to receive natural language responses. | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答
image
View & run on Rokha →
The phone book — and the kitchen — of the agentic world. Search 190k+ skills and MCP servers, then run them for real.