Advanced search

GitHub projects

956 entries

TRL

huggingface.co

TRL is a full stack library where we provide a set of tools to train transformer language models with Reinforcement Learning, from the Supervised Fine-tuning step (SFT), Reward Modeling step (RM) to the Proximal Policy Optimization (PPO) step.

Tool GitHub projects

Kimi-K2

huggingface.co

</details>

Tool GitHub projects

Moonlight-A3B

huggingface.co

Kimi-VL-A3B

Tool GitHub projects

InternLM2-1.8|7|20B

huggingface.co

InternLM-Math-7B|20B

Tool GitHub projects

StarCoder-1|3|7B

huggingface.co

StarCoder2-3|7|15B

Tool GitHub projects

RWKV-v4|5|6

huggingface.co

minicpm-2b-65d48bf958302b9fd25b698f)

Tool GitHub projects

MiniCPM-2B

huggingface.co

OmniLLM-12B

Tool GitHub projects

GLM-2|6|10|13|70B

huggingface.co

CogVLM2-19B

Tool GitHub projects

Nemotron-4-340B

huggingface.co

</details>

Tool GitHub projects

Baichuan-7|13B

huggingface.co

Baichuan2-7|13B

Tool GitHub projects

Yi-VL-6B|34B

huggingface.co

</details>

Tool GitHub projects

Yi-34B

huggingface.co

Yi1.5-6|9|34B

Tool GitHub projects

Command R-35B

huggingface.co

</details>

Tool GitHub projects

OLMo-7B

huggingface.co

</details>

Tool GitHub projects

OpenELM-1.1|3B

huggingface.co

</details>

Tool GitHub projects

Llama 1-7|13|33|65B

ai.facebook.com

OPT-1.3|6.7|13|30|66B

Tool GitHub projects

Qwen-1.8B|7B|14B|72B

huggingface.co

Qwen1.5-0.5B|1.8B|4B|7B|14B|32B|72B|110B|MoE-A2.7B

Tool GitHub projects

DeepSeek-VL-1.3|7B

huggingface.co

DeepSeek-MoE-16B

Tool GitHub projects

DeepSeek-Math-7B

huggingface.co

DeepSeek-Coder-1.3|6.7|7|33B

Tool GitHub projects

WHOOPS!

whoops-benchmark.github.io

a benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations.

Tool GitHub projects

VisualWebArena

jykoh.com

a benchmark designed to assess the performance of multimodal web agents on realistic visually grounded tasks.

Tool GitHub projects

TAT-QA

nextplusplus.github.io

a large-scale question-answering benchmark focused on real-world financial data, integrating both tabular and textual information.

Tool GitHub projects

SuperLim

lab.kb.se

a Swedish language understanding benchmark that evaluates natural language processing (NLP) models on various tasks such as argumentation analysis, semantic similarity, and textual entailment.

Tool GitHub projects

PubMedQA

pubmedqa.github.io

a biomedical question-answering benchmark designed for answering research-related questions using PubMed abstracts.

Tool GitHub projects

MMToM-QA

chuanyangjin.com

a multimodal question-answering benchmark designed to evaluate AI models' cognitive ability to understand human beliefs and goals.

Tool GitHub projects

MMedBench

henrychur.github.io

a benchmark that evaluates large language models' ability to answer medical questions across multiple languages.

Tool GitHub projects

MathEval

matheval.ai

a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nearly 30,000 math problems.

Tool GitHub projects

M3CoT

lightchen233.github.io

a benchmark that evaluates large language models on a variety of multimodal reasoning tasks, including language, natural and social sciences, physical and social commonsense, temporal reasoning, algebra, and geometry.

Tool GitHub projects

LLMEval

llmeval.com

focuses on understanding how these models perform in various scenarios and analyzing results from an interpretability perspective.

Tool GitHub projects

LawBench

lawbench.opencompass.org.cn

a benchmark designed to evaluate large language models in the legal domain.

Tool GitHub projects

InfiBench

infi-coder.github.io

a benchmark designed to evaluate large language models (LLMs) specifically in their ability to answer real-world coding-related questions.

Tool GitHub projects

FELM

hkust-nlp.github.io

a meta-benchmark that evaluates how well factuality evaluators assess the outputs of large language models (LLMs).

Tool GitHub projects

CompMix

qa.mpi-inf.mpg.de

a benchmark evaluating QA methods that operate over a mixture of heterogeneous input sources (KB, text, tables, infoboxes).

Tool GitHub projects

CompassRank

rank.opencompass.org.cn

CompassRank is dedicated to exploring the most advanced language and visual models, offering a comprehensive, objective, and neutral evaluation reference for the industry and research.

Tool GitHub projects

AlpacaEval

tatsu-lab.github.io

An Automatic Evaluator for Instruction-following Language Models using Nous benchmark suite.

Tool GitHub projects

Open LLM Leaderboard

huggingface.co

aims to track, rank, and evaluate LLMs and chatbots as they are released.

Tool GitHub projects

Chatbot Arena Leaderboard

huggingface.co

a benchmark platform for large language models (LLMs) that features anonymous, randomized battles in a crowdsourced manner.

Tool GitHub projects

OpenAI o3-mini

openai.com

Pushing the frontier of cost-effective reasoning.

Tool GitHub projects

AWS Samples

github.com

A collection of code samples from AWS which can be adapted for use with Claude. Note that some samples may require modification to work optimally with Claude.

Tool GitHub projects

Ensemble Learning Paper

cs.nju.edu.cn

Ensembling models with R, Ensembling Regression Models in R, Intro to Ensembles in R

Tool GitHub projects

Submit a site to the catalog

Just send the link — we will work out the rest.

We will review what you send and add it to the catalog if it fits.

Not sure how to implement it? We can help

Tell us about your task — we will pick the tools and suggest where to start.

0 / 5000
Verification code

Fields marked with an asterisk are required. Your data is used only to reply.