スキル監査レポート
evaluating-code-models 監査レポート.
Evaluates code generation models across HumanEval, MBPP, MultiPL-E, and 15+ benchmarks with pass@k metrics. Use when benchmarking code models, comparing coding abilities, testing multi-language support, or measuring code generation quality. Industry standard from BigCode Project used by HuggingFace leaderboards.
OpenAgentSkill Trust Score
OpenAgentSkill Trust Score
Trust Score は、インストール前に候補に入れる安全性を Agent が判断する助けになります。
GitHub 採用度
情報76
GitHub スター 825
スター/フォーク活動
情報71
スター 825、フォーク 176; 現在のメタデータでは Issue 活動を利用できません
最近のメンテナンス
合格88
最終プッシュから 1 か月
ライセンスの明確さ
合格86
MIT
README/SKILL.md の完全性
合格94
メタデータには十分な利用・ワークフロー文脈があります
依存関係/ランタイムのリスク
警告44
command execution surface, credential or environment access
インストール可否
合格92
npx skills add sangrokjung/claude-forge --skill evaluating-code-models
インストールコマンドの安全性
合格92
標準パッケージまたはランタイムのインストールパス
権限範囲
失敗36
secrets or environment access, shell or command execution
リポジトリ根拠
合格86
https://github.com/sangrokjung/claude-forge/tree/main/skills/evaluating-code-models
レビュー状況
合格88
AI レビューデータを利用できます
Agent 検証結果
情報54
Agent の成果データはまだありません
チェック
インストールと採用のレビュー
インストール経路
92
npx skills add sangrokjung/claude-forge --skill evaluating-code-models
リポジトリ
88
https://github.com/sangrokjung/claude-forge/tree/main/skills/evaluating-code-models
ライセンス
86
MIT
メンテナンス
88
最終プッシュから 1 か月
AI レビュー
88
Approved with no listed issues
README/SKILL.md の完全性
94
Usable description available
依存関係リスク
44
command execution surface, credential or environment access
インストールコマンドの安全性
92
標準パッケージまたはランタイムのインストールパス
権限範囲
36
secrets or environment access, shell or command execution
スター/フォーク活動
71
スター 825、フォーク 176; 現在のメタデータでは Issue 活動を利用できません
採用度
88
GitHub スター 825
警告
- Dependency or permission surface needs review
- Permission surface may require sandboxing
- Quality score needs review
- Permission surface needs review: secrets or environment access, shell or command execution
- Dependency/runtime risk: command execution surface, credential or environment access
- Permission surface: secrets or environment access, shell or command execution
方法
このレポートは公開メタデータ、AI レビュー、リポジトリの鮮度、インストール準備、OpenAgentSkill イベント、品質スコア、信頼チェック、Agent セーフティゲートを統合します。完全なソースコード監査ではありません。
近い選択肢を比較
次に監査する関連スキル
Code Review
Review a branch or diff against repository standards and the originating spec in two independent analysis passes.
169K スター · 監査レポート
Appsmith
Platform to build admin panels, internal tools, and dashboards. Integrates with 25+ databases and any API.
41K スター · 監査レポート
Implement
Implement work from an approved spec or ticket set, run focused and full tests, invoke code review, and commit the result to the current branch.
176K スター · 監査レポート