No visual example yet
Explore the skillParseBench
run-llama
ParseBench - A Document Parsing Benchmark for AI Agents
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
52 Skills
Results: 52
No visual example yet
Explore the skillrun-llama
ParseBench - A Document Parsing Benchmark for AI Agents
No visual example yet
Explore the skillmajiayu000
Provides guidance for automatically evolving and optimizing AI agents across any domain using LLM-driven evolution algorithms. Use when building self-improving agents, o…
No visual example yet
Explore the skillTIGER-AI-Lab
Open-source benchmark for browser AI agents on daily tasks.
No visual example yet
Explore the skillalibaba
An Alibaba open-source multi-language benchmark for evaluating LLMs in repository-level automatic code review, featuring an AI-assisted and expert-verified dataset.
No visual example yet
Explore the skillrzhub
GateMem: a benchmark and evaluation toolkit for memory governance in multi-principal shared-memory LLM agents.
No visual example yet
Explore the skillyzhao062
A Python library for anomaly detection across tabular, time series, graph, text, and image data. 60+ detectors, benchmark-backed ADEngine orchestration, and an agentic w…
No visual example yet
Explore the skillFareedKhan-dev
35 production-grade agentic AI architectures (Reflexion, LATS, GraphRAG, MemGPT, Voyager, BrowserAgent, ...) — a Python library and runnable textbook with multi-provider…
No visual example yet
Explore the skillBarca0412
入门资料整理:1.多因子股票量化框架开源教程 2.学界和业界的经典资料收录 3.AI + 金融的相关工作,包括LLM, Agent, benchmark(evaluation), etc.
No visual example yet
Explore the skillDoorman11991
AI coding agent optimized for small LLMs. 87% benchmark with 4B-active model.
No visual example yet
Explore the skillTHUDM
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
No visual example yet
Explore the skillDietrichGebert
Show ponytail measured impact as a scoreboard: less code, less cost, more speed, from the benchmark medians. One-shot display.
No visual example yet
Explore the skillevo-hq
turns your codebase into an autoresearch loop — discovers what to measure, instruments the benchmark, then runs tree search with parallel subagents.
No visual example yet
Explore the skillK-Dense-AI
Molecular ML with diverse featurizers and pre-built datasets. Use for property prediction (ADMET, toxicity) with traditional ML or GNNs when you want extensive featuriza…
No visual example yet
Explore the skillTopdu
OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Pars…
No visual example yet
Explore the skillK-Dense-AI
Autonomously improve a real artifact (code, training recipe, agent harness, data pipeline, prompt) against an objective and an evaluator, using Hypothesis Tree Refinemen…
No visual example yet
Explore the skillanthropics
Framework for building competitive landscape decks — market positioning, competitor deep-dives, comparative analysis, strategic synthesis. Use when the user asks for a c…