FlexGen
Running large language models on a single GPU for throughput-oriented scenarios. (Archived)
Running large language models on a single GPU for throughput-oriented scenarios. (Archived)
Kubernetes operator for LLM inference with pluggable runtimes (llama.cpp, PersonaPlex/Moshi, generic), multi-GPU sharding, NVIDIA CUDA and Apple Silicon Metal support, and GGUF/MLX/SafeTensors model formats.
OpenAI compatible API for LLMs and embeddings (LLaMA, Vicuna, ChatGLM and many others)
Open-source iOS/Android app running LLMs on-device via llama.cpp. Voice (Whisper), vision, image gen, tool calling — fully offline.
OpenAI-compatible LLM inference server for Apple Silicon using MLX. 2-4x faster than Ollama with tool calling and prompt caching.
Large Language Model Text Generation Inference
is a 4x faster and low-memory usage drop-in cli replacement that supports word-level timestamps and VAD filter
Port of OpenAI's Whisper model in C/C++
Real-time inference for Stable Diffusion - 0.88s latency. Covers AITemplate, nvFuser, TensorRT, FlashAttention. (Archived)
Build multimodal AI services via cloud native technologies · Model Serving · Generative AI · Neural Search · Cloud Native
A machine learning model serving framework with dynamic batching and pipelined stages, provides an easy-to-use Python interface.
A flexible, high-performance serving system for machine learning models.
Serve, optimize and scale PyTorch models in production (Archived)
The Triton Inference Server provides an optimized cloud and edge inferencing solution.
Serverless LLM apps on Production with Jina AI Cloud (Archived)
FastAPI framework to build production-grade LLM applications
LLMs on Ray - RayLLM (Archived)
Replace OpenAI GPT with another LLM in your app by changing a single line of code. Xinference gives you the freedom to use any LLM you need. With Xinference, you're empowered to run inference with any open-source language models, speech recognition models, and multimodal models, whether in the clou
Deploy and scale machine learning models on Kubernetes. Built for LLMs, embeddings, and speech-to-text.
A Kubernetes operator that simplifies serving and tuning large AI models (e.g. Falcon or phi-3) using container images and GPU auto-provisioning. Includes an OpenAI-compatible server for inference and preset configurations for popular runtimes such as vLLM and transformers.
CLI tool for running coding agents inside hardware-isolated microVMs with snapshot isolation, egress control, and MCP authorization.
Open-source confidential AI framework for secure LLM deployment with data privacy, providing hardware-enforced isolation using Intel TDX and NVIDIA Confidential Computing.
A Python Machine Learning Pentesting Toolbox for Adversarial Attacks. Works with LLMs, DNNs, and other machine learning algorithms.
"Batteries included" logging solution for your Azure OpenAI instance.
Regression testing for AI agents. Snapshot behavior, detect tool-call and output regressions, with golden-baseline diffing and LLM-as-judge scoring. Supports LangGraph, CrewAI, OpenAI, Claude, and any HTTP API.
Evaluate, monitor, analyze, and improve machine learning and generative models from pre-production to production. Ship more ML and LLMs into production, and monitor ML and LLM metrics like hallucination, PII, and toxicity.
Deterministic verification protocol for LLM outputs using 8 formal verification engines (SymPy, Z3, AST, SQLGlot). Prevents hallucinations through mathematical proofs rather than statistical methods.
Always know what to expect from your data.
Open source LLM observability platform. One line of code to monitor, evaluate, and experiment with features like prompt management, agent tracing, and evaluations.
OpenTelemetry-based observability and monitoring for LLM and agents workflows.
The open standard for data logging
Lightweight Go CLI that tracks AI API quota usage across 7 providers (Anthropic, OpenAI, GitHub Copilot, MiniMax, and more). Background daemon, <50MB RAM, zero telemetry, SQLite storage.
CLI tool for debugging and benchmarking RAG retrieval. EXPLAIN ANALYZE for your retrieval layer.
Open-source AI tracing framework built on OpenTelemetry for deep observability across agentic and LLM workflows.
Production-grade SDK for observability, automated evaluations and prompt management with sub-100ms guardrails for LLM/agent workflows.
Visualizes RAG knowledge gaps and "blind spots" using 2D UMAP clustering and density detection.