DataComPy
A library to compare Pandas, Polars, and Spark data frames. It provides stats and lets users adjust for match accuracy.
A library to compare Pandas, Polars, and Spark data frames. It provides stats and lets users adjust for match accuracy.
Drop-in AsyncOpenAI replacement that transparently batches requests via the Batch API for cheaper LLM inference.
Lightweight ML utility for automated training, evaluation, and prediction with CLI and Python API support.
Multidimensional cluster generation in Python.
Evaluate, trace, test, and ship LLM applications across your dev and production lifecycles.
A python machine learning library created to combine powefull data analasys features with tensors and machine learning components, while maintaining support for other libraries.
Auto-Magical CI/CD to streamline your AI workload. Experiment Management, Data Management, Pipeline, Orchestration, Scheduling & Serving in one MLOps/LLMOps solution.
The best-in-class MLOps platform with experiment tracking, model production monitoring, a model registry, and data lineage from training straight through to production.
A bounded controller for production ML under distribution shift — detects drift, learns from delayed labels, and takes the smallest safe steering step to defer unnecessary retrains.
Frouros is an open source Python library for drift detection in machine learning systems.
Algorithmic Trading with Machine Learning.
AutoML for Image, Text, Tabular, Time-Series, and MultiModal Data.
The standard data-centric AI package for data quality and machine learning with messy, real-world data and labels.
A Python library for Bayesian Evidential Learning (BEL) in order to estimate the uncertainty of a prediction.
A tutorial to help machine learning researchers to automatically obtain optimized machine learning models with the optimal learning performance on any specific task.
Free automated data & feature enrichment library for machine learning - automatically searches through thousands of ready-to-use features from public and community shared data sources and enriches your training dataset with only the accuracy improving features.
Skrub is a Python library that eases preprocessing and feature engineering for machine learning on dataframes.
Eurybia monitors data and model drift over time and securizes model deployment with data validation.
Shapash is a Python library that provides several types of visualization that display explicit labels that everyone can understand.
Validation & testing of machine learning models and data during model development, deployment, and production. This includes checks and suites related to various types of issues, such as model performance, data integrity, distribution mismatches, and more.
Optuna is an automatic hyperparameter optimization software framework, particularly designed for machine learning.
Streamlit is an framework to create beautiful data apps in hours, not weeks.
Interactive reports to analyze machine learning models during validation or production monitoring.
An AutoML package for hyperparameters tuning using evolutionary algorithms, with built-in callbacks, plotting, remote logging and more.
An AutoML framework for the automated design of composite modelling pipelines. It can handle classification, regression, and time series forecasting tasks on different types of data (including multi-modal datasets).
A framework for general purpose online machine learning.
Backprop makes it simple to use, finetune, and deploy state-of-the-art ML models.
An easy-to-use, Python-based feature store. Optimized for time-series data.
Multidimensional synthetic data generation in Python.
Fastest unstructured dataset management for TensorFlow/PyTorch. Stream & version-control data. Store even petabyte-scale data in a single numpy-like array on the cloud accessible on any machine. Visit activeloop.ai for more info.
A Python library for quickly creating and sharing demos of models. Debug models interactively in your browser, get feedback from collaborators, and generate public links without deploying anything.
Python-based meta-heuristic optimization techniques.
A Python-inspired implementation of the Optimum-Path Forest classifier.
A unified framework for machine learning with time series
Peer-to-peer network of data owners and data scientists who can collectively train AI models using PySyft
A Python library for secure and private Deep Learning built on PyTorch and TensorFlow.
Scalable deep learning training platform, including integrated support for distributed training, hyperparameter tuning, experiment tracking, and model management.
A fast Evolution Strategy implementation in Python.
An Automated Machine Learning (AutoML) python package for tabular data. It can handle: Binary Classification, MultiClass Classification and Regression. It provides explanations and markdown reports.
A simple, but essential Bayesian optimization package, written in Python.
A Pytorch based framework that breaks down machine learning problems into smaller blocks that can be glued together seamlessly with objective to build predictive models with one line of code.
A machine learning framework for multi-output/multi-label and stream data.
High-level wrapper built on the top of Pytorch which supports vision, text, tabular data and collaborative filtering.
High-level utils for PyTorch DL & RL research. It was developed with a focus on reproducibility, fast experimentation and code/ideas reusing. Being able to research/develop something new, rather than write another regular train loop.
JAX is Autograd and XLA, brought together for high-performance machine learning research.
A comparative framework for multimodal recommender systems with a focus on models leveraging auxiliary data.
A framework providing the right abstractions to ease research, development, and deployment of your ML pipelines.
Reference implementations of ML models written in numpy