Skip to content
#

skill-evaluation

Here are 27 public repositories matching this topic...

oh-my-knowledge

Evaluation framework for LLM knowledge inputs — prompts, RAG corpora, skills, agent workflows. Fix the model, vary the artifact. Built-in statistical rigor: bootstrap CI, Krippendorff α, length-debias, saturation curves.

  • Updated Jul 20, 2026
  • TypeScript

面向 AI Agent 的评测与上线决策平台。两条评测线:① LLM-as-judge 多裁判评分 + 失败归因 + 多轮 trial + 判官对齐(TPR/TNR);② skill 评测/优化(接 SkillOpt,确定性判分)。读轨迹辅助灰度发版(React · FastAPI · Postgres)

  • Updated Jul 15, 2026
  • TypeScript

Improve this page

Add a description, image, and links to the skill-evaluation topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the skill-evaluation topic, visit your repo's landing page and select "manage topics."

Learn more