暂未收录效果图
查看技能说明Chinese Llm Benchmark
jeinlee1991
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat…
OPENAGENTSKILL / DIRECTORY
为下一项任务找到合适的技能。探索适用于 Codex、Claude Code、Cursor 等 Agent 的工具。
13 Skills
搜索结果: 13
暂未收录效果图
查看技能说明jeinlee1991
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat…
暂未收录效果图
查看技能说明zilliztech
暂未收录效果图
查看技能说明hyperledger-caliper
A blockchain benchmark framework to measure performance of multiple blockchain solutions https://wiki.hyperledger.org/display/caliper
暂未收录效果图
查看技能说明Ayanami0730
暂未收录效果图
查看技能说明alibaba
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
暂未收录效果图
查看技能说明rzhub
GateMem: a benchmark and evaluation toolkit for memory governance in multi-principal shared-memory LLM agents.
暂未收录效果图
查看技能说明FareedKhan-dev
35 production-grade agentic AI architectures (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, ...) — a Python library and runnable textbook with multi-provider…
暂未收录效果图
查看技能说明Doorman11991
AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.
暂未收录效果图
查看技能说明THUDM
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
暂未收录效果图
查看技能说明K-Dense-AI
Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinemen…
暂未收录效果图
查看技能说明EmergenceAI
Emergence World: A world designed to reveal what no benchmark can: emergent intelligence.
暂未收录效果图
查看技能说明google-research
Suite of human-collected datasets and a multi-task continuous control benchmark for open vocabulary visuolinguomotor learning.
暂未收录效果图
查看技能说明texttron
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)