Megatron-DeepSpeed
DeepSpeed version of NVIDIA's Megatron-LM that adds additional support for several features such as MoE model training, Curriculum Learning, 3D Parallelism, and others.
DeepSpeed version of NVIDIA's Megatron-LM that adds additional support for several features such as MoE model training, Curriculum Learning, 3D Parallelism, and others.
A Native-PyTorch Library for LLM Fine-tuning.
An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models.
veRL is a flexible and efficient RL framework for LLMs.
Efficient Training for Big Models.
Mesh TensorFlow: Model Parallelism Made Easier.
A simple, performant and scalable Jax LLM!
An implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.
A library for accelerating Transformer model training on NVIDIA GPUs.
An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & RingAttention & RFT).
Open-source framework for fine-tuning and evaluating LLMs. It simplifies the process of experimenting with different training configurations and makes it easy to reproduce and share results, supporting features like LoRA, QLoRA, DeepSpeed, PEFT, and multi-GPU setups.
Nvidia Framework for LLM Inference
NVIDIA Framework for LLM Inference(Transitioned to TensorRT-LLM)
To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
A more memory-efficient rewrite of the HF transformers implementation of Llama for use with quantized weights.
Blazingly fast LLM inference.
Run LLMs and batch jobs on any cloud. Get maximum cost savings, highest GPU availability, and managed execution -- all with a simple interface.
Fine-tune, serve, deploy, and monitor any open-source LLMs in production. Used in production at BentoML for LLMs-based applications.
MII makes low-latency and high-throughput inference, similar to vLLM powered by DeepSpeed.
Inference for text-embeddings in Rust, HFOIL Licence.
Inference for text-embeddings in Python
A high-throughput and low-latency inference and serving framework for LLMs and VLs
Efficient Triton Kernels for LLM Training.
A distributed implementation of llama.cpp that lets you run 70B-level LLMs on your everyday devices.
Easily deploy any LLM on a VM with minimal configuration, using Ansible.
DSPy: The framework for programming—not prompting—foundation models.
A popular Python/JavaScript library for chaining sequences of language model prompts.
A Python library for augmenting LLM apps with data.
Comprehensive set of tools for working with local LLMs for various tasks.
Lightweight alternative to LangChain for composing LLMs
Seamlessly integrate LLMs as Python functions
Use ChatGPT On Wechat via wechaty
Test your prompts. Evaluate and compare LLM outputs, catch regressions, and improve prompt quality.
Easily build, version, evaluate and deploy your LLM-powered apps.
a chat interface crafted with llama.cpp for running Alpaca models. No API keys, entirely self-hosted!
Framework to create ChatGPT like bots over your dataset.