PuzzleTide Agent Evals
v0.1.1Use this skill when the user wants verifiable reasoning tasks to benchmark or test an LLM or agent — reproducible puzzle task sets (sudoku, word search) with objective, by-construction grading. No answer key to trust: answers are verified against the rules and the grid.
0· 0·0 当前·0 累计
运行时依赖
无特殊依赖
安装命令
点击复制官方npx clawhub@latest install puzzletide-agent-evals
镜像加速npx clawhub@latest install puzzletide-agent-evals --registry https://cn.longxiaskill.com 镜像可用