技能目录

为 AI Agent 发现可复用技能。

按任务搜索真实的 GitHub 技能,并在使用前查看 Stars、信任、审计、分类和安装路径。

每个推荐都保留与其仓库、审计和安装路径的明确关联。

搜索结果: chunking-algorithm

英文目录

Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.

15K
Stars
86/100
信任
分类: document-processing审计

🌟100+ 原创 LLM / RL 原理图📚,《大模型算法》作者巨献!💥(100+ LLM/RL Algorithm Maps )

4.5K
Stars
74/100
信任
分类: ml-automation审计

Open-source GEO content engineering and multi-site distribution system with AI tasks, RAG/semantic chunking, analytics, GEOFlow Agent and WordPress target publishing.

2.6K
Stars
85/100
信任
分类: growth-marketing审计

Source code of PyGAD, a Python 3 library for building the genetic algorithm and training machine learning algorithms (Keras & PyTorch).

2.2K
Stars
85/100
信任
分类: ml-automation审计

This repository provides an advanced Retrieval-Augmented Generation (RAG) solution for complex question answering. It uses sophisticated graph based algorithm to handle the tasks.

1.6K
Stars
84/100
信任
分类: rag-knowledge审计

Framework for quantitative trading. Complete framework for development, backtesting, and deploying automated trading algorithms and trading bots.

1.3K
Stars
79/100
信任
分类: finance审计

A modern Anki custom scheduling based on Free Spaced Repetition Scheduler algorithm

4.0K
Stars
83/100
信任
分类: ml-automation审计

Enhances Tesseract OCR output using LLMs (local or API) for error correction, smart chunking, and markdown formatting of scanned PDFs

2.9K
Stars
73/100
信任
分类: document-processing审计

人工智能学习路线图,整理近200个实战案例与项目,免费提供配套教材,零基础入门,就业实战!包括:Python,数学,机器学习,数据分析,深度学习,计算机视觉,自然语言处理,PyTorch tensorflow machine-learning,deep-learning data-analysis data-mining mathematics data-science artificial-intelligence python tensorflow tensorflow2 caffe keras pytorch algorithm numpy pandas matplotlib seaborn nlp cv等热门领域

13K
Stars
71/100
信任
分类: data-analysis审计

This skill should be used for time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search. Use when working with temporal data, sequential patterns, or time-indexed observations requiring specialized algorithms beyond standard ML approaches. Particularly suited for univariate and multivariate time series analysis with scikit-learn compatible APIs.

34K
Stars
80/100
信任
分类: research审计

Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for large-scale datasets.

34K
Stars
77/100
信任
分类: data-analysis审计

🚀 efficient approximate nearest neighbor search algorithm collections library written in Rust 🦀 .

2.7K
Stars
76/100
信任
分类: rag-knowledge审计