Direktori skill

Temukan skill yang dapat digunakan kembali untuk AI agents.

Cari skill GitHub nyata berdasarkan tugas lalu periksa stars, trust, audit, kategori, dan jalur pemasangan sebelum digunakan.

Setiap rekomendasi tetap terhubung dengan repositori, audit, dan jalur pemasangannya.

Hasil pencarian: ingestion

Direktori bahasa Inggris

๐Ÿฆ› CHONK docs with Chonkie โœจ โ€” The lightweight ingestion library for fast, efficient and robust RAG pipelines

4.5K
Stars
85/100
Kepercayaan
Kategori: dataAudit

LakeSoul is an end-to-end, realtime and cloud native Lakehouse framework with fast data ingestion, concurrent update and incremental data analytics on cloud storages for both BI and AI applications.

3.2K
Stars
83/100
Kepercayaan
Kategori: data-analysisAudit

Specify a github or local repo, github pull request, arXiv or Sci-Hub paper, Youtube transcript or documentation URL on the web and scrape into a text file and clipboard for easier LLM ingestion

2.0K
Stars
80/100
Kepercayaan
Kategori: researchAudit

Fastest end-to-end CSV ingestion for Ruby (with C acceleration). SmarterCSV auto-detects formats, applies smart defaults, and returns Rails-ready hashes for seamless use with ActiveRecord, Sidekiq, parallel jobs, and S3 pipelines โ€” even for messy user-uploaded real-world data.

1.5K
Stars
83/100
Kepercayaan
Kategori: data-analysisAudit

OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). โšก Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.

1.4K
Stars
83/100
Kepercayaan
Kategori: data-analysisAudit

Highly Performant, Modular, Memory Safe and Production-ready Inference, Ingestion and Indexing built in Rust ๐Ÿฆ€

1.3K
Stars
84/100
Kepercayaan
Kategori: dataAudit

Installable agent skills package for ClickHouse and chdb best practices, schema design, query optimization, and data ingestion patterns.

493
Stars
75/100
Kepercayaan
Kategori: dataAudit

File Parser optimised for LLM Ingestion with no loss ๐Ÿง  Parse PDFs, Docx, PPTx in a format that is ideal for LLMs.

7.4K
Stars
74/100
Kepercayaan
Kategori: document-processingAudit
Arc71

High-performance analytical database. 19.9M records/sec ingestion, 8.4M+ rows/sec queries. Ingestion, compaction, SQL, retention, continuous queries โ€” one binary. Open Parquet on your storage. S3/Azure native. Air-gap ready. No vendor lock-in. AGPL-3.0.

609
Stars
71/100
Kepercayaan
Kategori: devopsAudit

SiteWhere is an industrial strength open-source application enablement platform for the Internet of Things (IoT). It provides a multi-tenant microservice-based infrastructure that includes device/asset management, data ingestion, big-data storage, and integration through a modern, scalable architecture. SiteWhere provides REST APIs for all system functionality. SiteWhere provides SDKs for many common device platforms including Android, iOS, Arduino, and any Java-capable platform such as Raspberry Pi rapidly accelerating the speed of innovation.

1.0K
Stars
68/100
Kepercayaan
Kategori: devopsAudit

๐Ÿ“ˆ A scalable, production-ready data pipeline for real-time streaming & batch processing, integrating Kafka, Spark, Airflow, AWS, Kubernetes, and MLflow. Supports end-to-end data ingestion, transformation, storage, monitoring, and AI/ML serving with CI/CD automation using Terraform & GitHub Actions.

124
Stars
68/100
Kepercayaan
Kategori: data-analysisAudit

For a detached wiki ingester only. The main agent never loads this. Defines orientation protocol, page format, role guardrail, validator contract, manifest protocol, and deterministic completion contract. Loaded by the provider-aware runtime worker.

101
Stars
63/100
Kepercayaan
Kategori: automationAudit