Inference
xorbitsai
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all throug…
OPENAGENTSKILL / DIRECTORY
Find a skill for your next task. Explore tools for Codex, Claude Code, Cursor and more.
Find a skill for your next task. Preview examples where available.
1–16 / 78
Results: 78
xorbitsai
Swap GPT for any LLM by changing a single line of code. Xinference lets you run open-source, speech, and multimodal models on cloud, on-prem, or your laptop — all throug…
dusty-nv
Hello AI World guide to deploying deep-learning inference networks and deep vision primitives with TensorRT and NVIDIA Jetson.
huggingface
A blazing fast inference solution for text embeddings models
triton-inference-server
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
inference-sh
inference.sh Agent skills for using our API to give your agents access to hundreds of apps and other agents
deepspeedai
DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
bentoml
The easiest way to serve AI apps and models - Build Model Inference APIs, Job queues, LLM apps, Multi-model pipelines, and more!
py-why
DoWhy is a Python library for causal inference that supports explicit modeling and testing of causal assumptions. DoWhy is based on a unified language for causal inferen…
uber
Uplift modeling and causal inference with machine learning algorithms
Michael-A-Kuykendall
⚡ Pure-Rust WebGPU inference engine — OpenAI-API compatible, GGUF native, runs on any GPU. No Python. No llama.cpp. Single binary.
vllm-project
A framework for efficient model inference with omni-modality models
py-why
ALICE (Automated Learning and Intelligence for Causation and Economics) is a Microsoft Research project aimed at applying Artificial Intelligence concepts to economic de…
hao-ai-lab
A unified inference and post-training framework for accelerated video generation.
llm-d
Achieve state of the art inference performance with modern accelerators on Kubernetes
spiceai
A portable accelerated SQL query, search, and LLM-inference engine, written in Rust, for data-grounded AI apps and agents.
dstackai
Vendor-agnostic orchestration for training, inference and agentic workloads across NVIDIA, AMD, TPU, and Tenstorrent on clouds, Kubernetes, and bare metal.