Evaluate
๐ค Evaluate: A library for easily evaluating machine learning models and datasets.
๊ณต๊ธ ์์ฐ ํ๋กํ
์ฝ๋ฉ ๋ฐ ๊ฐ๋ฐ Agent
์ฝ๋ ๋ฆฌ๋ทฐ, ์ ์ฅ์ ๋ถ์, ํ ์คํธ, CI, GitHub, DevOps ๋ฐ ๊ฐ๋ฐ ์ํฌํ๋ก์ฉ ์คํฌ์ ๋๋ค.
์๋๋ฆฌ์ค
์ฝ๋ฉ Agent
์ ์ฅ์๋ฅผ ์ดํดํ๊ณ ์ฝ๋๋ฅผ ์์ ํ๋ฉฐ Pull Request๋ฅผ ๊ฒํ ํ ์ ์๋ ์ฝ๋ฉ Agent๊ฐ ํ์ํฉ๋๋ค.
Agent ์ ํฉ๋
Claude Code + CLI + Codex
Codex, Claude Code, Cursor, CLI ๋๋ ๋ง์ถคํ Agent์ ์ ํฉํฉ๋๋ค.
์ค์น
์ค๋น๋จ
npx skills add huggingface/evaluate
์ ์ง๋ณด์
ํ์ฑ
๋ง์ง๋ง ํธ์ ํ 3๊ฐ์
์ํ
๊ฒํ ํ์
Financial research output is not financial advice; require human review before any live investment decision
GitHub ํ์ง
2.5K
99/100 ํ์ง ยท 90/100 ์ ๋ขฐ
์ปค๋ฒ๋ฆฌ์ง ํ๊ทธ
๊ฒํ ๋ฉ๋ชจ
Financial research output is not financial advice; require human review before any live investment decision ยท Financial research output is not financial advice; require human review before any live investment decision.
Agent ์ฑํ ์ค์ฝ์ด์นด๋
์ ๋ขฐ, ๊ฐ์ฌ, ์ค์น ์ค๋น ์ํ๋ฅผ ํ๋์ ํ์ธํ์ธ์
์ด ์ ์๋ ๊ณต๊ฐ ์ ์ฅ์ ๋ฉํ๋ฐ์ดํฐ, OpenAgentSkill ๊ฒํ ์ ํธ, ์ ์ง๋ณด์ ์ต์ ์ฑ, ์ค์น ์ค๋น ์ํ๋ฅผ ๊ฒฐํฉํฉ๋๋ค. ํ๋ณด ์ ์ ์ ํธ์ผ ๋ฟ, ์ฌ๋์ ๊ฒํ ๋ฅผ ๋์ฒดํ์ง ์์ต๋๋ค.
ํ์ง
์ฐ์๊ฐํ ์ฑํ ๋ฐ ์ ์ง๋ณด์ ์ ํธ๋ฅผ ๊ฐ์ถ ์ ๋ขฐ๋ ๋์ ์ถ์ฒ์ ๋๋ค.
์ ๋ขฐ
๊ฒํ ํ ์ค์น์ข์ ํ๋ณด ์ ํธ์ด์ง๋ง Agent๋ ์คํ ์ ์ ๊ฐ์ฌ ๋ฉ๋ชจ, ์ค์น ์ ์ฑ ๋ฐ ๊ฒฐ๊ณผ ๊ทผ๊ฑฐ๋ฅผ ๊ฒํ ํด์ผ ํฉ๋๋ค.
๊ฐ์ฌ
๊ฒํ ํ์์ค์น ์ค๋น ์ํ, ๋ณด์ ๋ฉํ๋ฐ์ดํฐ, ์ ์ง๋ณด์ ๋ฐ ์ฑํ ์ํ์ ๋ํ ๊ธฐ๊ณ ํ๋ ํ ๊ฒํ ์ ๋๋ค.
OpenAgentSkill ์ ๋ขฐ ์ ์ v5
์ค์น ์ ์ฌ๋ ๊ฒํ
์ฌ๋ ๊ฒํ ๋๋ ์๋๋ฐ์ค ๊ฒ์ฆ ํ ์ฐ์ ํ๋ณด๋ก ์ฌ์ฉํ์ธ์.
์คํ
GitHub ์คํ 2.5K
์ ์ฅ์ ํ๋
์คํ 2.5K, ํฌํฌ 321
์ ์ง๋ณด์
๋ง์ง๋ง ํธ์ ํ 3๊ฐ์
๋ผ์ด์ ์ค
Apache-2.0
์ค์น
npx skills add huggingface/evaluate
์ค์น ์์ ์ฑ
ํ์ค ํจํค์ง ๋๋ ๋ฐํ์ ์ค์น ๊ฒฝ๋ก
๊ถํ ๋ฒ์
ํ์ผ ์์คํ ๋๋ ๋ฌธ์ ์ ๊ทผ
Agent ๊ฒฐ๊ณผ
์์ง Agent ๊ฒฐ๊ณผ ๋ฐ์ดํฐ๊ฐ ์์ต๋๋ค
๋ฌธ์
README/SKILL.md ๋งฅ๋ฝ์ด ์ถฉ๋ถํฉ๋๋ค
์ํ ์์ฝ
๋ฎ์ ๋ฉํ๋ฐ์ดํฐ ์ํ
- Financial research output is not financial advice; require human review before any live investment decision.
์ค์น ์ค๋น ์ํ
์ค์น ๊ฒฝ๋ก ์ฌ์ฉ ๊ฐ๋ฅ
- ์ค์น ๊ฒฝ๋ก๋ฅผ ์ฌ์ฉํ ์ ์์ต๋๋ค
- ์ ์ฅ์ ๊ทผ๊ฑฐ๋ฅผ ์ฌ์ฉํ ์ ์์ต๋๋ค
- ๋ผ์ด์ ์ค๊ฐ ๋ช ์๋์์ต๋๋ค
- ์์ง Agent ๊ฒ์ฆ ๊ฒฐ๊ณผ ๊ทผ๊ฑฐ๊ฐ ์์ต๋๋ค
Agent ์ฝ๊ธฐ์ฉ ๋ฉํ๋ฐ์ดํฐ
์ด ์คํฌ์ ๊ธฐ๊ณ ํ๋ ํ ์์ฌ๊ฒฐ์ ๋ฐ์ดํฐ.
์ด ๋ธ๋ก ๋๋ ํฌํจ๋ JSON์ ์ฌ์ฉํด Agent๊ฐ ์ด ์คํฌ์ ์ค์นํ ์ง, ๋์์ ๊ณ ๋ฅผ์ง, ๋จผ์ ์ฌ๋์ ๊ฒํ ๋ฅผ ์์ฒญํ ์ง ํ๋จํ ์ ์์ต๋๋ค.
์ ํฉํ ์์
- ์ฝ๋ฉ Agent ์ํฌํ๋ก
- Claude Code ํ
- GitHub ์ฑํ ์ ํธ๋ฅผ ์ค์ํ๋ ํ
- Inspect source files
์ ํฉํ Agent
์ค์น ๊ฒฐ์
- ๋ช ๋ น์ด
- npx skills add huggingface/evaluate
- ์ ์ฑ
- ๊ฒํ
- ์ฌ๋ ๊ฒํ
- ์
์ ๋ขฐ์ ์ํ
- ์ ๋ขฐ
- 85/100
- ๊ฐ์ฌ
- 92/100
- ์ํ ์์ค
- ๊ฒํ ํ์
๊ฒฐ๊ณผ ๋ฃจํ
- ์๋ํฌ์ธํธ
- /api/agent/outcome
- ์ด๋ฒคํธ ID
- resolve
- ๊ฒฐ๊ณผ
- 5
์ค์น ๋ช ๋ น์ด
npx skills add huggingface/evaluate์ฌ์ฉํ์ง ๋ง์์ผ ํ ๊ฒฝ์ฐ
- ๋ฒค๋ ์ง์ SLA๊ฐ ํ์ํ ํ
- ๋ด๋ถ ๋ณด์ ๊ฒํ ๊ฐ ์๋ ๊ณ ๊ท์ ์ค์ ํ๊ฒฝ
- ํ์ฌ ๋ฉํ๋ฐ์ดํฐ์์ ์ฃผ์ ์ํ ์ ํธ๊ฐ ๋ฐ๊ฒฌ๋์ง ์์์ต๋๋ค
- Financial research output is not financial advice; require human review before any live investment decision
- Financial research output is not financial advice; require human review before any live investment decision.
Agent ์์ v2
76/100 ยท ์ค์น ์ ๊ฒํ
์ฌ์ฉ ๊ฐ๋ฅํ ํ๋ณด์ด์ง๋ง Agent๋ ์ค์น ์ ์ ๊ถํ๊ณผ ๊ฐ์ฌ ๋ฉ๋ชจ๋ฅผ ํ์ํด์ผ ํฉ๋๋ค.
์ค์ ์์ ๊ณต๊ฐ์ ์ค์นํ๊ธฐ ์ ์ ์ฌ๋์ ์น์ธ์ด ํ์ํฉ๋๋ค.
์ค๊ฐ
๋คํธ์ํฌ ์ ๊ทผ
Skill์ ์๊ฒฉ ํ์ด์ง, API, ์ ์ฅ์ ๋๋ ์ธ๋ถ ์๋น์ค์ ์ ๊ทผํ ์ ์์ต๋๋ค.
์ค๊ฐ
ํ์ผ ์์คํ ์ ๊ทผ
Skill์ ํ๋ก์ ํธ ํ์ผ, ๋ฌธ์, ์์ฑ ์ฐ์ถ๋ฌผ ๋๋ ๋ก์ปฌ ์์ ๊ณต๊ฐ ์ํ๋ฅผ ์ฝ๊ฑฐ๋ ์ธ ์ ์์ต๋๋ค.
- Financial research output is not financial advice; require human review before any live investment decision
์ค์น ๋์
Agent ์ํฌํ๋ก์ ์ด ์คํฌ ์ค์น
๊ณต๊ฐ ์ค์น ์๋ํฌ์ธํธ์์ ๋ช ๋ น์ด, ์์ ์ฒดํฌ๋ฆฌ์คํธ, ๋์ ํ๋กฌํํธ์ ์ ๊ท ๋งํฌ๋ฅผ ๊ฐ์ ธ์ต๋๋ค.
OpenAgentSkill CLI
Resolve policy, run the source installer safely, and report a verified install receipt.
$ npx --yes https://github.com/Leon-Drq/openagentskill/releases/download/cli-v0.2.1/openagentskill-0.2.1.tgz install huggingface-evaluateAgent ํด๊ฒฐ ๊ณํ
์ค์น ์ ์ Agent๊ฐ ์ ํฉ์ฑ์ ๊ฒ์ฆํ๊ฒ ํ์ธ์.
Resolve API๋ ์ต์ฐ์ ์คํฌ, ๋์, ์์ ์ ์ฑ , ๊ฐ์ฌ ๋ฉ๋ชจ, ์ค์น ๋์ ๋ฐ Agent๊ฐ ํ์ด์ง๋ฅผ ์คํฌ๋ํํ์ง ์๊ณ ์ฌ์ฉํ ์ ์๋ ํ๋กฌํํธ๋ฅผ ๋ฐํํฉ๋๋ค.
JSON ์ด๊ธฐ
/api/agent/resolve?task=Use%20Evaluate%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Resolve ํ ์คํธ
/api/agent/resolve?task=Use%20Evaluate%20for%20an%20agent%20workflow&agent=codex&max_risk=medium&format=text
์ค์น ํธ๋์คํ
/api/skills/huggingface-evaluate/install
Agent๊ฐ ํ์ธํ ํญ๋ชฉ
- Resolve API์์ ์์ ์ ํฉ๋์ ๋์์ ํ์ธํฉ๋๋ค.
- ๊ฐ์ฌ ์ ์, ์ ๋ขฐ ์ ์ ๋ฐ ์์ ์ ์ฑ ๊ฒฝ๊ณ ๋ฅผ ํ์ธํฉ๋๋ค.
- Codex, Claude Code, Cursor ๋๋ CLI์ ์ค์น ๋์ ํธํ์ฑ์ ํ์ธํฉ๋๋ค.
ํ๋กฌํํธ ๋ณต์ฌ
Task: Use Evaluate in this workspace.
Resolve first: https://www.openagentskill.com/api/agent/resolve?task=Use%20Evaluate%20for%20an%20agent%20workflow&agent=codex&max_risk=medium
Review install handoff: https://www.openagentskill.com/api/skills/huggingface-evaluate/install
Install command: npx skills add huggingface/evaluate
Before running it, summarize audit warnings, required permissions, and the fallback skill if install is risky.Agent ํธ๋์คํ
๋ ๋ค๋ฅธ ๋๋ ํฐ๋ฆฌ ํ์ด์ง ๋์ ์ค์น ๊ฒฝ๋ก๋ฅผ Agent์๊ฒ ์ ๊ณตํฉ๋๋ค.
๊ณต๊ฐ ์ค์น ์๋ํฌ์ธํธ์์ ๋ช ๋ น์ด, ์์ ์ฒดํฌ๋ฆฌ์คํธ, ๋์ ํ๋กฌํํธ์ ์ ๊ท ๋งํฌ๋ฅผ ๊ฐ์ ธ์ต๋๋ค.
์ค์น ํธ๋์คํ
/api/skills/huggingface-evaluate/install
LLM ํ ์คํธ ํ์
/api/skills/huggingface-evaluate/install?format=text
๋์ ์ฐพ๊ธฐ
/api/skills/search?q=Evaluate&limit=3
Agent ํ๋กฌํํธ
Use Evaluate for this task. Review https://www.openagentskill.com/api/skills/huggingface-evaluate/install, then install with: npx skills add huggingface/evaluateRegistry ๋ฉํ๋ฐ์ดํฐ
์๋ ์คํฌ ์ ํ์ ์ํ Agent ์ฝ๊ธฐ์ฉ ํ๋กํ.
Registry API๋ฅผ ํตํด ๋์ผํ ๊ฒฐ์ , ์ ๋ขฐ, ๊ฐ์ฌ, ์ฌ์ฉ ์ฌ๋ก, ์ค์น ์ ํธ๋ฅผ ์ ๊ณตํ๋ฏ๋ก Agent๊ฐ UI๋ฅผ ์คํฌ๋ํํ์ง ์๊ณ ๋ ์์๋ฅผ ๋งค๊ธธ ์ ์์ต๋๋ค.
Manifest
/api/registry/manifest/huggingface-evaluate
LLM ํ ์คํธ
/api/registry/manifest/huggingface-evaluate?format=text
์ค์น ๋ณ์นญ
/api/registry/install/huggingface-evaluate
์ถ์ฒ
/api/registry/recommend?task=Use%20Evaluate%20in%20an%20agent%20workflow&limit=3
Agent ์ ํฉ๋
์ฝ๋ฉ Agent
์ฌ์ฉ ์ฌ๋ก ํ๊ทธ
ํ๋ซํผ
Python, Machine Learning, Claude Code
๊ฐ์ฌ ๋ณด๊ณ ์
๊ฒํ ํ์ ยท 92/100
์ค์น ์ค๋น ์ํ, ๋ณด์ ๋ฉํ๋ฐ์ดํฐ, ์ ์ง๋ณด์ ๋ฐ ์ฑํ ์ํ์ ๋ํ ๊ธฐ๊ณ ํ๋ ํ ๊ฒํ ์ ๋๋ค.
Agent ๊ฒฐ์ ํจ๋
์ฝ๋ฉ Agent์ฉ ์ฐ์ ์ถ์ฒ
์ฐ์ ํ๋ณด๋ก ์ฌ์ฉํ๋, ์์ ์ Agent ํ๊ฒฝ์์ README์ ์ค์น ๊ฒฝ๋ก๋ฅผ ๊ฒ์ฆํ์ธ์.
์คํ ๋ด ์ญํ
์ฐ์ ์ถ์ฒ
์ฃผ์ ์ ํฉ๋
์ฝ๋ฉ Agent
์ ๋ขฐ ๋ผ๋ฒจ
ํ๋ก๋์ ์ค๋น ์๋ฃ
์ค์น ๊ฒฝ๋ก
๋ช ๋ น์ด ์ค๋น๋จ
์ฌ์ฉ ์์
- ์ฝ๋ฉ Agent ์ํฌํ๋ก
- Claude Code ํ
- GitHub ์ฑํ ์ ํธ๋ฅผ ์ค์ํ๋ ํ
๊ทผ๊ฑฐ
- GitHub ์คํ 2,455
- ์ต๊ทผ ์ ์ฅ์ ํ๋
- ์ค์น ๋ช ๋ น ๋๋ GitHub ์ ์ฅ์๋ฅผ ์ฌ์ฉํ ์ ์์ต๋๋ค
- ํ์ง ํ๋กํ 99/100
- OpenAgentSkill ์ํธ์์ฉ 10๊ฑด
๋จผ์ ๊ฒํ
- ํ์ฌ ๋ฉํ๋ฐ์ดํฐ์์ ์ฃผ์ ์ํ ์ ํธ๊ฐ ๋ฐ๊ฒฌ๋์ง ์์์ต๋๋ค
๊ตฌํ ๊ฒฝ๋ก
- 1์๋๋ฐ์ค Agent์ ์ค์นํ๊ณ ์ฝ๋ฉ Agent ์์ ์ ์ฒ์๋ถํฐ ๋๊น์ง ํ ๋ฒ ์คํํ์ธ์.
- 2Compare output quality, latency, and failure behavior against at least one alternative.
- 3Promote it into production only after reviewing repository permissions, license, and maintenance signals.
์ ๋ขฐ ํ๋กํ
๊ฒํ ํ ์ค์น
์ข์ ํ๋ณด ์ ํธ์ด์ง๋ง Agent๋ ์คํ ์ ์ ๊ฐ์ฌ ๋ฉ๋ชจ, ์ค์น ์ ์ฑ ๋ฐ ๊ฒฐ๊ณผ ๊ทผ๊ฑฐ๋ฅผ ๊ฒํ ํด์ผ ํฉ๋๋ค.
GitHub ์ฑํ๋
ํต๊ณผGitHub ์คํ 2.5K
์คํ/ํฌํฌ ํ๋
ํต๊ณผ์คํ 2.5K, ํฌํฌ 321; ํ์ฌ ๋ฉํ๋ฐ์ดํฐ์์ ์ด์ ํ๋์ ํ์ธํ ์ ์์ต๋๋ค
์ต๊ทผ ์ ์ง๋ณด์
ํต๊ณผ๋ง์ง๋ง ํธ์ ํ 3๊ฐ์
๋ผ์ด์ ์ค ๋ช ํ์ฑ
ํต๊ณผApache-2.0
๊ธ์ ์ ํธ
- ์๋ ๊ฒ์ฆ๋ ๋ฑ๋ก
- AI ๊ฒํ ์น์ธ๋จ
- ์ค์น ๊ฒฝ๋ก๋ฅผ ์ฌ์ฉํ ์ ์์ต๋๋ค
- ์ ์ฅ์ ๊ทผ๊ฑฐ๋ฅผ ์ฌ์ฉํ ์ ์์ต๋๋ค
- ์ต๊ทผ ์ ์ง๋ณด์๋ ์ ์ฅ์
- ์๋ฏธ ์๋ GitHub ์ฑํ ์ ํธ
- ์ค์น ๋ช ๋ น์์ ๋๋ ทํ ๊ณ ์ํ ํจํด์ด ๋ฐ๊ฒฌ๋์ง ์์์ต๋๋ค
- ๊ฒฐ๊ณผ ๋ฃจํ๋ ์ค๋น๋์์ง๋ง ์ฒซ ์ค์ Agent ์คํ์ด ํ์ํฉ๋๋ค
์ค์น ์ ๊ฒํ
- Financial research output is not financial advice; require human review before any live investment decision.
- ์์ง ์ค์ Agent ๊ฒฐ๊ณผ ๋ณด๊ณ ์๊ฐ ์์ต๋๋ค
- ๋ฌด์ธ ์ค์น ์ ์ ์ฌ๋ ๊ฒํ ๊ฐ ํ์ํฉ๋๋ค
๊ถ์ฅ ์์
์ฌ๋ ๊ฒํ ๋๋ ์๋๋ฐ์ค ๊ฒ์ฆ ํ ์ฐ์ ํ๋ณด๋ก ์ฌ์ฉํ์ธ์.
ํ์ง ํ๋กํ
์ฐ์ Agent ์ํฌํ๋ก์ฉ ํ๋ณด
๊ฐํ ์ฑํ ๋ฐ ์ ์ง๋ณด์ ์ ํธ๋ฅผ ๊ฐ์ถ ์ ๋ขฐ๋ ๋์ ์ถ์ฒ์ ๋๋ค.
์ํฌํ๋ก ์ ํฉ๋
์ด ์คํฌ์ ์ฌ์ฉํ ์๋๋ฆฌ์ค
Build and ship code
Coding agents
I need a coding agent that can understand a repository, edit code, and review pull requests.
Search private knowledge
RAG and knowledge
I need my agent to build a RAG workflow over documents and retrieve reliable context.
Automate repeated work
Workflow automation
I need my agent to automate a repeated workflow across tools and files.
์ํฌํ๋ก ์ ํฉ๋
์์ ํ ์ํฌํ๋ก์ ์ถ๊ฐ
Turn skills into distribution
Content growth agent
A workflow for turning newly indexed skills into SEO briefs, social drafts, comparison pages, and reusable publishing workflows.
Inspect, patch, and verify code
Coding review agent
A workflow for software agents that inspect repositories, review pull requests, generate tests, and turn findings into shippable patches.
Ingest, retrieve, and cite
RAG knowledge base
A workflow for document-heavy agents that ingest files, create searchable knowledge, retrieve relevant context, and answer with grounded sources.
๋์ ํ๋ณด
์ค์น ์ ๋น๊ต
์ด ์์ ์ ์ ํฉํ ์ ์๋ ์ ์ฌ ์คํฌ์ ๋๋ค.
Tensorflow
An Open Source Machine Learning Framework for Everyone
Transformers
๐ค Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
Pytorch
Tensors and Dynamic neural networks in Python with strong GPU acceleration
LLMs From Scratch
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
๊ฐ์
๐ค Evaluate: A library for easily evaluating machine learning models and datasets.
Imported by the skill-only GitHub discovery pipeline because it matches agent skill, automation, domain workflow, RAG, document-processing, data, finance, security, or developer-tool signals. Protocol-server projects are excluded from automated imports.
ํ๋ซํผ ํธํ์ฑ
๊ธฐ์ ์ธ๋ถ ์ฌํญ
- ๋ฒ์
- 1.0.0
- ๋ผ์ด์ ์ค
- Apache-2.0
- ์ต๊ทผ ์ ๋ฐ์ดํธ
- 2026๋ 8์ 18์ผ
- ๊ฒ์์ผ
- 2026๋ 6์ 20์ผ
ํ๋ ์์ํฌ ๋ฐ ๋๊ตฌ
๊ฒฐ์ ์ค๋ ์ท
์ฐ์ ์ถ์ฒ
GitHub ์คํ 2,455
๊ฐ์ฌ
์ค์น ๊ฒํ
์ค์น ๋ฐ ์ฑํ ๊ฒํ
- ๋ณด์
- 89/100
- ์ ์ง๋ณด์
- 88/100
- ์ค์น
- 92/100
Agent ๊ฒ์ฆ ์ฆ๊ฑฐ
Agent ๊ฒ์ฆ ์ฆ๊ฑฐ
Resolve, ๊ฒํ , ์ค์น ๋ฐ ํ ๋ฒ์ ์ ํ๋ ์คํ ํ ๊ฒฐ๊ณผ ๋ณด๊ณ ์์ ๋๋ค.
- ์ฑ๊ณต๋ฅ
- โ
- ์ต๊ทผ ์คํจ
- โ
- ๊ฒฐ๊ณผ
- 0
- ์ถ๋ ฅ ํ์ง
- โ
- ์คํจ
- 0
- ๊ด๋ จ ์์
- 0
- ์ค์น
- 0
- ์ํ ์ฐจ๋จ
- 0
- ์ค์ ํ์
- 0
- ํ๋ก๋์
- 0
์์ง Agent ๊ฒฐ๊ณผ ๋ฐ์ดํฐ๊ฐ ์์ต๋๋ค. ์ฒซ ์คํ์ /api/agent/outcome์ ํตํด ์ฑ๊ณต, ์ค์ ํ์, ์ํ ์ฐจ๋จ, ์คํจ ๋๋ ๋น๊ด๋ จ ๊ฒฐ๊ณผ๋ฅผ ๋ณด๊ณ ํ ์ ์์ต๋๋ค.
์ค์น
Agent ์ํฌํ๋ก์ ์ถ๊ฐ
๋ฌด๋ฃ ์คํ ์์ค. ํ๋ก๋์ Agent์ ์ค์นํ๊ธฐ ์ ์ ๋ณด๊ณ ์๋ฅผ ๊ฒํ ํ์ธ์.
์ฑ์ฅ ๋ฃจํ
๊ณต์ ํคํธ
Evaluate์ฉ ์๋๋ฆฌ์ค ๊ธฐ๋ฐ ์ด์์ ๋๋ค. X์ ์๋์ผ๋ก ๊ฒ์ํ ์ ์์ต๋๋ค.
For source-backed research, this is a skill worth shortlisting before another blank prompt. Evaluate: ๐ค Evaluate: A library for easily evaluating machine learning models and datasets. 2.5K stars https://www.openagentskill.com/skills/huggingface-evaluate?ref=x
์ ํ ์ฌํญ: ์ค์น ๋ช ๋ น์ด ํฌํจ๋ ๋ต๊ธ
Listing + install path for Evaluate: https://www.openagentskill.com/skills/huggingface-evaluate?ref=x Install: npx skills add huggingface/evaluate
๋ฑ๋ก ์ถ์ฒ
์ปค๋ฎค๋ํฐ ์์ธ
์ด ๋ฑ๋ก์ ๊ณต๊ฐ ์์ค์์ ์์ธ๋์์ผ๋ฉฐ ์ ์ง๋ณด์์ ์์ ๊ถ ์ฃผ์ฅ์ด ์น์ธ๋ ๋๊น์ง ๊ณต์์ผ๋ก ํ์๋์ง ์์ต๋๋ค.
- ์ ์์
- huggingface
- ์ถ์ฒ
- huggingface/evaluate
- ์์ธ ์ฃผ์ฒด
- OpenAgentSkill ์ปค๋ฎค๋ํฐ ์ธ๋ฑ์ค
๊ท์์ ๊ณต๊ฐ ์ ์ฅ์ ๋๋ ์ ์์ ํ๋กํ์ ์ฐ๊ฒฐ๋ฉ๋๋ค. ์ ์์๋ ๋ฑ๋ก์ ์ฃผ์ฅํ์ฌ ์์ ๊ถ ์ ํธ๋ฅผ ์ ๋ฐ์ดํธํ ์ ์์ต๋๋ค.
์ด ์คํฌ ์์ ๊ถ ์ฃผ์ฅ์์ ์ ์์ ๊ถ ์ฃผ์ฅ
์ด ์คํฌ ๋ฑ๋ก ์์ ๊ถ ์ฃผ์ฅ
์ด ์ปค๋ฎค๋ํฐ ์์ธ ๋ฑ๋ก์ huggingface์๊ฒ ๊ท์๋์ด ์์ง๋ง ์์ง ๊ณต์์ผ๋ก ํ์๋์ง ์์์ต๋๋ค. ์์ ๊ถ์ ์ฃผ์ฅํ๋ฉด ํ์ธ๋ ์์ ์ ์ ํธ๊ฐ ์ถ๊ฐ๋์ด ์ดํ ์ถ์, ์ค์น ๋ฐ ๊ฐ์ฌ ์ ๋ฐ์ดํธ๋ฅผ ๋ ์ ๋ขฐํ ์ ์์ต๋๋ค.
ํฌ๋ฆฌ์์ดํฐ ๋ฐฑ๋งํฌ ํคํธ
README์ ์ฆ๊ฑฐ ๋ฐฐ์ง ์ถ๊ฐ
๊ฐ๋ฐ์๊ฐ ์ ์ฅ์๋ฅผ ํ๊ฐํ๋ ์์น์ ์ ๊ท ๋ฑ๋ก, ํ์ฌ ์ ๋ขฐ ๋ฐ ๊ฐ์ฌ ์ ํธ, ์ค์ Agent-Proven ์ฆ๊ฑฐ๋ฅผ ํ์ํฉ๋๋ค.
[](https://www.openagentskill.com/skills/huggingface-evaluate)
[](https://www.openagentskill.com/skills/huggingface-evaluate)
[](https://www.openagentskill.com/skills/huggingface-evaluate/audit)
[](https://www.openagentskill.com/skills/huggingface-evaluate)์์ฑ์
huggingfaceโ
@huggingface
ํ๋ซํผ ์ ํฉ๋
์ํ ์ ํธ
- GitHub ์คํ
- 2.5K
- ํ์ง ์ ์
- 62/100
- ์ต๊ทผ GitHub ํธ์
- 2026๋ 5์ 26์ผ
- ํ๋ ์์ํฌ ํํธ
- 2
- OpenAgentSkill ์กฐํ์
- 7
- ์ค์น ๋ช ๋ น ๋ณต์ฌ
- 0
- ์ธ๋ถ ํด๋ฆญ
- 0
์ปค๋ฎค๋ํฐ ์ ํธ
์ด ์คํฌ์ด Agent ์ํฌํ๋ก์ ์ ์ฉํ์ง ์๋ ค ์ฃผ์ธ์. ์ง๊ณ๋ ํผ๋๋ฐฑ์ ์๊ฐ์ด ์ง๋ ์๋ก ์์๋ฅผ ๊ฐ์ ํฉ๋๋ค.
์ ๋ขฐ์ ์์
๊ฒํ ํ ์ค์น
- GitHub ์ฑํ๋GitHub ์คํ 2.5Kํต๊ณผ
- ์คํ/ํฌํฌ ํ๋์คํ 2.5K, ํฌํฌ 321; ํ์ฌ ๋ฉํ๋ฐ์ดํฐ์์ ์ด์ ํ๋์ ํ์ธํ ์ ์์ต๋๋คํต๊ณผ
- ์ต๊ทผ ์ ์ง๋ณด์๋ง์ง๋ง ํธ์ ํ 3๊ฐ์ํต๊ณผ
- ๋ผ์ด์ ์ค ๋ช ํ์ฑApache-2.0ํต๊ณผ
- README/SKILL.md ์์ฑ๋๋ฉํ๋ฐ์ดํฐ์ ์ถฉ๋ถํ ์ฌ์ฉ ๋ฐ ์ํฌํ๋ก ๋งฅ๋ฝ์ด ํฌํจ๋์ด ์์ต๋๋คํต๊ณผ
- ์์กด์ฑ/๋ฐํ์ ์ํ๊ณต๊ฐ ๋ฉํ๋ฐ์ดํฐ์ ์ฃผ์ ์์กด์ฑ ์ํ ํํธ๊ฐ ์์ต๋๋คํต๊ณผ
๊ด๋ จ ์คํฌ
Tensorflow
An Open Source Machine Learning Framework for Everyone
195.7K ์คํTransformers
๐ค Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
161.6K ์คํPytorch
Tensors and Dynamic neural networks in Python with strong GPU acceleration
100.8K ์คํLLMs From Scratch
Implement a ChatGPT-like LLM in PyTorch from scratch, step by step
97.3K ์คํ