MLEM
Version and deploy your ML models following GitOps principles
Version and deploy your ML models following GitOps principles
Ready to use deeplearning docker images.
Aqueduct enables you to easily define, run, and manage AI & ML tasks on any cloud infrastructure.
Ambrosia helps you clean up your LLM datasets using other LLMs.
Open-source LLM evaluation and red teaming framework. Test prompts, models, agents, and RAG pipelines. Run adversarial attacks (jailbreaks, prompt injection) and integrate security testing into CI/CD.
Open-source CLI security scanner for agentic workflows. Scans your workflow’s source code, detects vulnerabilities, and generates an interactive visualization along with a detailed security report. Supports LangGraph, CrewAI, n8n, OpenAI Agents, and more.
Open-source runtime security scanner for AI agents. Detects prompt injection, jailbreak, PII leakage, memory poisoning, and tool misuse. Zero deps, MIT licensed.
Visual AI agent workflow automation platform with local LLM integration. Build intelligent workflows using drag-and-drop, no cloud required.
Open source Kubernetes-style control plane for deploying AI agents as distributed microservices, with built-in service discovery, durable workflows, and observability.
Chrome extension that uses local LLMs to assist with writing and drafting responses based on the context of your open tabs.
Godot 4.x asset that enables NPCs to interact with players using local LLMs for structured, offline-first learning conversations in games.
Curated list of top Hugging Face models for NLP, vision, and audio tasks with demos and benchmarks.
A curated collection of battle-tested tools, frameworks, and best practices for building, scaling, and monitoring production-grade Retrieval-Augmented Generation (RAG) systems. Covers frameworks, vector databases, retrieval & reranking, evaluation, observability, deployment, and security.
Open-source SDK for running LLMs and multimodal models on-device across iOS, Android, and cross-platform apps.
agentic AI operating system (h9y.ai) that replaces brittle/fragmented automations with long-lived, self-improving systems. Open-source, self-hosted/cloud, visual workflow, omni-channel, decentralized, extensible.
A VS Code extension for viewing and exploring large machine learning datasets (CSV, JSON, Parquet, etc.) directly within the editor without VS Code crashing in a clean UI.
A VS Code extension to view Weights & Biases experiments, logs, and artifacts within the IDE, eliminating the need to switch to the web UI and keeping data private.
DeepSeek's official setup guides for integrating its models with coding agents including Claude Code, Codex, Cline, OpenCode, and Pi.
An open-source framework and registry for evaluating language models and systems.
An open-source, workflow-first terminal coding agent with subagents, skills, sandboxed tools, and multiple model providers.
Production-ready AI agent development kit for Rust with model-agnostic design (Gemini, OpenAI, Anthropic), multiple agent types (LLM, Graph, Workflow), MCP support, and built-in telemetry.
Agent framework for chatting with data, turning natural language into SQL, transformation pipelines and visualizations. Outputs are declarative specs that can be inspected, edited, reopened in a notebook or composed into a dashboard.
AI crypto trading framework using LightGBM + XGBoost ensemble with 72 ML features. 70.9% walk-forward validated accuracy on out-of-sample data. Supports Bybit and Binance. MIT licensed, available on PyPI.
Local AI agent for generating publication-ready scientific papers with real arXiv citations, IMRaD structure, and tribunal scoring. Runs 100% offline via Ollama with 4B-9B models. MIT licensed. HuggingFace
Open-source LLM and agent evaluation framework with 50+ metrics, LLM-as-Judge augmentation, and guardrail scanners (jailbreak, PII, prompt-injection). Useful for scoring RAG outputs, agent trajectories, and function-calling behavior in data-science workflows.
Open-source Python library and MCP server to benchmark document chunking strategies for RAG, score retrieval quality, and recommend configurations for a corpus.
Daily-updated skill and CLI for deterministic retrieval across arXiv, PubMed/PMC, and supported US policy corpora.
A weekly data project aimed at the R ecosystem.
DataCamp Cheatsheets Cheatsheets for data science.
Machine Learning, Data Science and Deep Learning with Python
Realtime deployment Tutorial on Python time-series model deployment.
A straightforward method for training your LLM, from downloading data to generating text.
Open Source Society University
A hands-on course to train and deploy a serverless API that predicts crypto prices.
Coursera Big Data Specialization
Data Science Degree @ Berkeley