Large Model Visual Question Answering Skill | 大模型视觉问答技能

On the Rokha Registry · clawhub · 0 Rokha runs · 1.8K downloads

Conducts open-ended Q&A on image content based on computer vision and large language models, supporting any questions to receive natural language responses. | 大模型视觉问答(VQA)技能,基于计算机视觉和大语言模型对图片内容进行开放式问答,支持任意提问得到自然语言回答

image

View & run on Rokha →

The phone book — and the kitchen — of the agentic world. Search 190k+ skills and MCP servers, then run them for real.