Type

Selection — 3 881

FlexGen

github.com

Running large language models on a single GPU for throughput-oriented scenarios. (Archived)

Repository GitHub projects

LLMKube

github.com

Kubernetes operator for LLM inference with pluggable runtimes (llama.cpp, PersonaPlex/Moshi, generic), multi-GPU sharding, NVIDIA CUDA and Apple Silicon Metal support, and GGUF/MLX/SafeTensors model formats.

Repository GitHub projects

Modelz-LLM

github.com

OpenAI compatible API for LLMs and embeddings (LLaMA, Vicuna, ChatGLM and many others)

Repository GitHub projects

Off Grid

github.com

Open-source iOS/Android app running LLMs on-device via llama.cpp. Voice (Whisper), vision, image gen, tool calling — fully offline.

Repository GitHub projects

Rapid-MLX

github.com

OpenAI-compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching.

Repository GitHub projects

whisper-ctranslate2

github.com

is a 4x faster and low-memory usage drop-in cli replacement that supports word-level timestamps and VAD filter

Repository GitHub projects

whisper.cpp

github.com

Port of OpenAI's Whisper model in C/C++

Repository GitHub projects

x-stable-diffusion

github.com

Real-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. (Archived)

Repository GitHub projects

Jina

github.com

Build multimodal AI services via cloud native technologies · Model Serving · Generative AI · Neural Search · Cloud Native

Repository GitHub projects

Mosec

github.com

A machine learning model serving framework with dynamic batching and pipelined stages, provides an easy-to-use Python interface.

Repository GitHub projects

TFServing

github.com

A flexible, high-performance serving system for machine learning models.

Repository GitHub projects

Torchserve

github.com

Serve, optimize and scale PyTorch models in production (Archived)

Repository GitHub projects

The Triton Inference Server provides an optimized cloud and edge inferencing solution.

Repository GitHub projects

langchain-serve

github.com

Serverless LLM apps on Production with Jina AI Cloud (Archived)

Repository GitHub projects

lanarky

github.com

FastAPI framework to build production-grade LLM applications

Repository GitHub projects

ray-llm

github.com

LLMs on Ray - RayLLM (Archived)

Repository GitHub projects

Xinference

github.com

Replace OpenAI GPT with another LLM in your app by changing a single line of code. Xinference gives you the freedom to use any LLM you need. With Xinference, you're empowered to run inference with any open-source language models, speech recognition models, and multimodal models, whether in the clou

Repository GitHub projects

KubeAI

github.com

Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.

Repository GitHub projects

Kaito

github.com

A Kubernetes operator that simplifies serving and tuning large AI models (e.g. Falcon or phi-3) using container images and GPU auto-provisioning. Includes an OpenAI-compatible server for inference and preset configurations for popular runtimes such as vLLM and transformers.

Repository GitHub projects

brood-box

github.com

CLI tool for running coding agents inside hardware-isolated microVMs with snapshot isolation, egress control, and MCP authorization.

Repository GitHub projects

dstack

github.com

Open-source confidential AI framework for secure LLM deployment with data privacy, providing hardware-enforced isolation using Intel TDX and NVIDIA Confidential Computing.

Repository GitHub projects

Plexiglass

github.com

A Python Machine Learning Pentesting Toolbox for Adversarial Attacks. Works with LLMs, DNNs, and other machine learning algorithms.

Repository GitHub projects

Azure OpenAI Logger

github.com

"Batteries included" logging solution for your Azure OpenAI instance.

Repository GitHub projects

EvalView

github.com

Regression testing for AI agents. Snapshot behavior, detect tool-call and output regressions, with golden-baseline diffing and LLM-as-judge scoring. Supports LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.

Repository GitHub projects

Fiddler AI

github.com

Evaluate, monitor, analyze, and improve machine learning and generative models from pre-production to production. Ship more ML and LLMs into production, and monitor ML and LLM metrics like hallucination, PII, and toxicity.

Repository GitHub projects

QWED

github.com

Deterministic verification protocol for LLM outputs using 8 formal verification engines (SymPy, Z3, AST, SQLGlot). Prevents hallucinations through mathematical proofs rather than statistical methods.

Repository GitHub projects

Great Expectations

github.com

Always know what to expect from your data.

Repository GitHub projects

Helicone

github.com

Open source LLM observability platform. One line of code to monitor, evaluate, and experiment with features like prompt management, agent tracing, and evaluations.

Repository GitHub projects

OpenTelemetry-based observability and monitoring for LLM and agents workflows.

Repository GitHub projects

whylogs

github.com

The open standard for data logging

Repository GitHub projects

onWatch

github.com

Lightweight Go CLI that tracks AI API quota usage across 7 providers (Anthropic, OpenAI, GitHub Copilot, MiniMax, and more). Background daemon, <50MB RAM, zero telemetry, SQLite storage.

Repository GitHub projects

RagTune

github.com

CLI tool for debugging and benchmarking RAG retrieval. EXPLAIN ANALYZE for your retrieval layer.

Repository GitHub projects

traceAI

github.com

Open-source AI tracing framework built on OpenTelemetry for deep observability across agentic and LLM workflows.

Repository GitHub projects

Future AGI

github.com

Production-grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows.

Repository GitHub projects

semantic-coverage

github.com

Visualizes RAG knowledge gaps and "blind spots" using 2D UMAP clustering and density detection.

Repository GitHub projects

Submit a site to the catalog

Just send the link — we will work out the rest.

We will review what you send and add it to the catalog if it fits.

Not sure how to implement it? We can help

Tell us about your task — we will pick the tools and suggest where to start.

0 / 5000
Verification code

Fields marked with an asterisk are required. Your data is used only to reply.