No visual example yet
Explore the skillChinese Llm Benchmark
jeinlee1991
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
9 Skills
Results: 9
No visual example yet
Explore the skilljeinlee1991
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat…
No visual example yet
Explore the skillembeddings-benchmark
MTEB: Massive Text Embedding Benchmark
No visual example yet
Explore the skillbeir-cellar
A Heterogeneous Benchmark for Information Retrieval. Easy to use, evaluate your models across 15+ diverse IR datasets.
No visual example yet
Explore the skillmajiayu000
Provides guidance for automatically evolving and optimizing AI agents across any domain using LLM-driven evolution algorithms. Use when building self-improving agents, o…
No visual example yet
Explore the skillrzhub
GateMem: a benchmark and evaluation toolkit for memory governance in multi-principal shared-memory LLM agents.
No visual example yet
Explore the skillFareedKhan-dev
35 production-grade agentic AI architectures (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, ...) — a Python library and runnable textbook with multi-provider…
No visual example yet
Explore the skillTHUDM
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
No visual example yet
Explore the skillszilard
A minimal benchmark for scalability, speed and accuracy of commonly used open source implementations (R packages, Python scikit-learn, H2O, xgboost, Spark MLlib etc.) of…
No visual example yet
Explore the skillDataScienceUIBK
🔥 Rankify: A Comprehensive Python Toolkit for Retrieval, Re-Ranking, and Retrieval-Augmented Generation 🔥. Our toolkit integrates 40 pre-retrieved benchmark datasets a…