Расширенный поиск

GitHub-проекты

961 записей

TGI

huggingface.co

a toolkit for deploying and serving Large Language Models (LLMs).

Инструмент GitHub-проекты

TRL

huggingface.co

TRL is a full stack library where we provide a set of tools to train transformer language models with Reinforcement Learning, from the Supervised Fine-tuning step (SFT), Reward Modeling step (RM) to the Proximal Policy Optimization (PPO) step.

Инструмент GitHub-проекты

Kimi-K2

huggingface.co

</details>

Инструмент GitHub-проекты

Moonlight-A3B

huggingface.co

Kimi-VL-A3B

Инструмент GitHub-проекты

InternLM2-1.8|7|20B

huggingface.co

InternLM-Math-7B|20B

Инструмент GitHub-проекты

StarCoder-1|3|7B

huggingface.co

StarCoder2-3|7|15B

Инструмент GitHub-проекты

RWKV-v4|5|6

huggingface.co

minicpm-2b-65d48bf958302b9fd25b698f)

Инструмент GitHub-проекты

MiniCPM-2B

huggingface.co

OmniLLM-12B

Инструмент GitHub-проекты

GLM-2|6|10|13|70B

huggingface.co

CogVLM2-19B

Инструмент GitHub-проекты

Nemotron-4-340B

huggingface.co

</details>

Инструмент GitHub-проекты

Baichuan-7|13B

huggingface.co

Baichuan2-7|13B

Инструмент GitHub-проекты

Yi-VL-6B|34B

huggingface.co

</details>

Инструмент GitHub-проекты

Yi-34B

huggingface.co

Yi1.5-6|9|34B

Инструмент GitHub-проекты

Command R-35B

huggingface.co

</details>

Инструмент GitHub-проекты

OLMo-7B

huggingface.co

</details>

Инструмент GitHub-проекты

OpenELM-1.1|3B

huggingface.co

</details>

Инструмент GitHub-проекты

Llama 1-7|13|33|65B

ai.facebook.com

OPT-1.3|6.7|13|30|66B

Инструмент GitHub-проекты

Qwen-1.8B|7B|14B|72B

huggingface.co

Qwen1.5-0.5B|1.8B|4B|7B|14B|32B|72B|110B|MoE-A2.7B

Инструмент GitHub-проекты

DeepSeek-VL-1.3|7B

huggingface.co

DeepSeek-MoE-16B

Инструмент GitHub-проекты

DeepSeek-Math-7B

huggingface.co

DeepSeek-Coder-1.3|6.7|7|33B

Инструмент GitHub-проекты

WHOOPS!

whoops-benchmark.github.io

a benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations.

Инструмент GitHub-проекты

VisualWebArena

jykoh.com

a benchmark designed to assess the performance of multimodal web agents on realistic visually grounded tasks.

Инструмент GitHub-проекты

TAT-QA

nextplusplus.github.io

a large-scale question-answering benchmark focused on real-world financial data, integrating both tabular and textual information.

Инструмент GitHub-проекты

SuperLim

lab.kb.se

a Swedish language understanding benchmark that evaluates natural language processing (NLP) models on various tasks such as argumentation analysis, semantic similarity, and textual entailment.

Инструмент GitHub-проекты

PubMedQA

pubmedqa.github.io

a biomedical question-answering benchmark designed for answering research-related questions using PubMed abstracts.

Инструмент GitHub-проекты

MMToM-QA

chuanyangjin.com

a multimodal question-answering benchmark designed to evaluate AI models' cognitive ability to understand human beliefs and goals.

Инструмент GitHub-проекты

MMedBench

henrychur.github.io

a benchmark that evaluates large language models' ability to answer medical questions across multiple languages.

Инструмент GitHub-проекты

MathEval

matheval.ai

a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nearly 30,000 math problems.

Инструмент GitHub-проекты

M3CoT

lightchen233.github.io

a benchmark that evaluates large language models on a variety of multimodal reasoning tasks, including language, natural and social sciences, physical and social commonsense, temporal reasoning, algebra, and geometry.

Инструмент GitHub-проекты

LLMEval

llmeval.com

focuses on understanding how these models perform in various scenarios and analyzing results from an interpretability perspective.

Инструмент GitHub-проекты

LawBench

lawbench.opencompass.org.cn

a benchmark designed to evaluate large language models in the legal domain.

Инструмент GitHub-проекты

InfiBench

infi-coder.github.io

a benchmark designed to evaluate large language models (LLMs) specifically in their ability to answer real-world coding-related questions.

Инструмент GitHub-проекты

FELM

hkust-nlp.github.io

a meta-benchmark that evaluates how well factuality evaluators assess the outputs of large language models (LLMs).

Инструмент GitHub-проекты

CompMix

qa.mpi-inf.mpg.de

a benchmark evaluating QA methods that operate over a mixture of heterogeneous input sources (KB, text, tables, infoboxes).

Инструмент GitHub-проекты

CompassRank

rank.opencompass.org.cn

CompassRank is dedicated to exploring the most advanced language and visual models, offering a comprehensive, objective, and neutral evaluation reference for the industry and research.

Инструмент GitHub-проекты

AlpacaEval

tatsu-lab.github.io

An Automatic Evaluator for Instruction-following Language Models using Nous benchmark suite.

Инструмент GitHub-проекты

Open LLM Leaderboard

huggingface.co

aims to track, rank, and evaluate LLMs and chatbots as they are released.

Инструмент GitHub-проекты

Chatbot Arena Leaderboard

huggingface.co

a benchmark platform for large language models (LLMs) that features anonymous, randomized battles in a crowdsourced manner.

Инструмент GitHub-проекты

OpenAI o3-mini

openai.com

Pushing the frontier of cost-effective reasoning.

Инструмент GitHub-проекты

AWS Samples

github.com

A collection of code samples from AWS which can be adapted for use with Claude. Note that some samples may require modification to work optimally with Claude.

Инструмент GitHub-проекты

Research Papers 1

mlg.eng.cam.ac.uk

, 2, 3

Инструмент GitHub-проекты

Ensemble Learning Paper

cs.nju.edu.cn

Ensembling models with R, Ensembling Regression Models in R, Intro to Ensembles in R

Инструмент GitHub-проекты

Предложить сайт в каталог

Пришлите ссылку — остальное мы выясним сами.

Мы рассмотрим, что вы прислали, и добавим в каталог, если подойдёт.

Не знаете, как внедрить? Мы поможем

Расскажите про задачу — подберём инструменты и подскажем, с чего начать.

0 / 5000
Проверочный код

Поля со звёздочкой обязательны. Данные используются только для ответа.