No visual example yet
Explore the skillChinese Llm Benchmark
jeinlee1991
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
13 Skills
Results: 13
No visual example yet
Explore the skilljeinlee1991
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat…
No visual example yet
Explore the skillzilliztech
Benchmark for vector databases.
No visual example yet
Explore the skillhyperledger-caliper
A blockchain benchmark framework to measure performance of multiple blockchain solutions https://wiki.hyperledger.org/display/caliper
No visual example yet
Explore the skillAyanami0730
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
No visual example yet
Explore the skillalibaba
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
No visual example yet
Explore the skillrzhub
GateMem: a benchmark and evaluation toolkit for memory governance in multi-principal shared-memory LLM agents.
No visual example yet
Explore the skillFareedKhan-dev
35 production-grade agentic AI architectures (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, ...) — a Python library and runnable textbook with multi-provider…
No visual example yet
Explore the skillDoorman11991
AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.
No visual example yet
Explore the skillTHUDM
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
No visual example yet
Explore the skillK-Dense-AI
Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinemen…
No visual example yet
Explore the skillEmergenceAI
Emergence World: A world designed to reveal what no benchmark can: emergent intelligence.
No visual example yet
Explore the skillgoogle-research
Suite of human-collected datasets and a multi-task continuous control benchmark for open vocabulary visuolinguomotor learning.
No visual example yet
Explore the skilltexttron
BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent (ACL 2026 Main)