dsh-vision-toolkit

技能包多媒体与生成中文文档MIT
](https://x.com/anion_ex) ](https://github.com/Anionex/dsh-vision-toolkit/releases/tag/v0.1.7) ](tests) ](LICENSE) ](package.json) ](runtime/requirements.lock) ](cordis.patch.yml)

功能特性

  • Act on coordinates instead of parsing prose: grounding and detection return original-image pixel boxes, while every model-visible result remains structured text or JSON.
  • Deliver files, not temporary output: crop, trace, OCR, pixel diff, foreground extraction, and HTML rendering produce described Artifacts that the Web client can preview, download, or open locally.
  • Close the visual verification loop: local HTML rendering and pixel-diff ranking support reference → implementation → screenshot → measured iteration without a model-native image channel.
  • Use the same bundle in Web and Headless profiles: Web adds cards, previews, Settings, and health actions; Headless receives the same tool semantics and complete structured results.
  • DeepSeek Harness with a Web or Headless profile and pnpm available to dsh plugin.
  • Python 3.11 or newer. Managed mode creates an isolated environment, so users do not install the upstream CLI or Python packages manually.

安装命令

dsh plugin --profile web add @anionex/dsh-vision-toolkit

项目预览

项目简介

让纯文本模型更好地做视觉任务的DeepSeek Harness插件:带意图的图片问答、长截图 OCR、UI 还原等|DeepSeek Harness-native integration for agent-vision-toolkit: image Q&A, long-screenshot OCR, UI restoration, grounding, pixel diff, Artifacts, and Web UI.

话题标签

#agent-skills#agent-vision-toolkit#computer-vision#deepseek#deepseek-harness#dsh#dsh-plugin#gui-automation#ocr#plugin#python#screenshot-testing
← 返回 技能包 列表