Directorio de skills

Descubre skills reutilizables para AI agents.

Busca skills reales de GitHub por tarea y revisa stars, confianza, auditoría, categoría y ruta de instalación antes de utilizarlos.

Cada recomendación conserva un vínculo claro con su repositorio, auditoría y ruta de instalación.

Resultados de búsqueda: chunking-algorithm

Directorio en inglés

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15K
Stars
86/100
Confianza
Categoría: document-processingAuditoría

🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm Maps )

4.5K
Stars
74/100
Confianza
Categoría: ml-automationAuditoría

Open-source GEO content engineering and multi-site distribution system with AI tasks, RAG/semantic chunking, analytics, GEOFlow Agent and WordPress target publishing.

2.6K
Stars
85/100
Confianza
Categoría: growth-marketingAuditoría

Source code of PyGAD, a Python 3 library for building the genetic algorithm and training machine learning algorithms (Keras & PyTorch).

2.2K
Stars
85/100
Confianza
Categoría: ml-automationAuditoría

This repository provides an advanced Retrieval-Augmented Generation (RAG) solution for complex question answering. It uses sophisticated graph based algorithm to handle the tasks.

1.6K
Stars
84/100
Confianza
Categoría: rag-knowledgeAuditoría

Framework for quantitative trading. Complete framework for development, backtesting, and deploying automated trading algorithms and trading bots.

1.3K
Stars
79/100
Confianza
Categoría: financeAuditoría

A modern Anki custom scheduling based on Free Spaced Repetition Scheduler algorithm

4.0K
Stars
83/100
Confianza
Categoría: ml-automationAuditoría

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

2.9K
Stars
73/100
Confianza
Categoría: document-processingAuditoría

人工智能学习路线图,整理近200个实战案例与项目,免费提供配套教材,零基础入门,就业实战!包括:Python,数学,机器学习,数据分析,深度学习,计算机视觉,自然语言处理,PyTorch tensorflow machine-learning,deep-learning data-analysis data-mining mathematics data-science artificial-intelligence python tensorflow tensorflow2 caffe keras pytorch algorithm numpy pandas matplotlib seaborn nlp cv等热门领域

13K
Stars
71/100
Confianza
Categoría: data-analysisAuditoría

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.

34K
Stars
80/100
Confianza
Categoría: researchAuditoría

Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.

34K
Stars
77/100
Confianza
Categoría: data-analysisAuditoría

🚀 efficient approximate nearest neighbor search algorithm collections library written in Rust 🦀 .

2.7K
Stars
76/100
Confianza
Categoría: rag-knowledgeAuditoría