No visual example yet
Explore the skillChinese Llm Benchmark
jeinlee1991
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat…
OPENAGENTSKILL / DIRECTORY
Finde den passenden Skill für deine nächste Aufgabe mit Codex, Claude Code, Cursor und mehr.
13 Skills
Ergebnisse: 13
No visual example yet
Explore the skilljeinlee1991
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat…
No visual example yet
Explore the skillK-Dense-AI
Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinemen…
No visual example yet
Explore the skillzilliztech
Benchmark for vector databases.
No visual example yet
Explore the skillhyperledger-caliper
A blockchain benchmark framework to measure performance of multiple blockchain solutions https://wiki.hyperledger.org/display/caliper
No visual example yet
Explore the skillAyanami0730
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
No visual example yet
Explore the skillalibaba
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
No visual example yet
Explore the skillrzhub
GateMem: a benchmark and evaluation toolkit for memory governance in multi-principal shared-memory LLM agents.
No visual example yet
Explore the skillFareedKhan-dev
35 production-grade agentic AI architectures (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, ...) — a Python library and runnable textbook with multi-provider…
No visual example yet
Explore the skillDoorman11991
AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.
No visual example yet
Explore the skillTHUDM
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
No visual example yet
Explore the skillEmergenceAI
Emergence World: A world designed to reveal what no benchmark can: emergent intelligence.
No visual example yet
Explore the skillgoogle-research
Suite of human-collected datasets and a multi-task continuous control benchmark for open vocabulary visuolinguomotor learning.
No visual example yet
Explore the skilltexttron
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)