TGI
a toolkit for deploying and serving Large Language Models (LLMs).
a toolkit for deploying and serving Large Language Models (LLMs).
TRL is a full stack library where we provide a set of tools to train transformer language models with Reinforcement Learning, from the Supervised Fine-tuning step (SFT), Reward Modeling step (RM) to the Proximal Policy Optimization (PPO) step.
</details>
Kimi-VL-A3B
InternVL-2|6|14|26
InternLM-Math-7B|20B
StarCoder2-3|7|15B
minicpm-2b-65d48bf958302b9fd25b698f)
OmniLLM-12B
CogVLM2-19B
</details>
Baichuan2-7|13B
</details>
Yi1.5-6|9|34B
</details>
</details>
</details>
OPT-1.3|6.7|13|30|66B
Qwen1.5-0.5B|1.8B|4B|7B|14B|32B|72B|110B|MoE-A2.7B
DeepSeek-MoE-16B
DeepSeek-Coder-1.3|6.7|7|33B
a benchmark dataset testing AI's ability to reason about visual commonsense through images that defy normal expectations.
a benchmark designed to assess the performance of multimodal web agents on realistic visually grounded tasks.
a large-scale question-answering benchmark focused on real-world financial data, integrating both tabular and textual information.
a Swedish language understanding benchmark that evaluates natural language processing (NLP) models on various tasks such as argumentation analysis, semantic similarity, and textual entailment.
a biomedical question-answering benchmark designed for answering research-related questions using PubMed abstracts.
a multimodal question-answering benchmark designed to evaluate AI models' cognitive ability to understand human beliefs and goals.
a benchmark that evaluates large language models' ability to answer medical questions across multiple languages.
a comprehensive benchmarking platform designed to evaluate large models' mathematical abilities across 20 fields and nearly 30,000 math problems.
a benchmark that evaluates large language models on a variety of multimodal reasoning tasks, including language, natural and social sciences, physical and social commonsense, temporal reasoning, algebra, and geometry.
focuses on understanding how these models perform in various scenarios and analyzing results from an interpretability perspective.
a benchmark designed to evaluate large language models in the legal domain.
a benchmark designed to evaluate large language models (LLMs) specifically in their ability to answer real-world coding-related questions.
a meta-benchmark that evaluates how well factuality evaluators assess the outputs of large language models (LLMs).
a benchmark evaluating QA methods that operate over a mixture of heterogeneous input sources (KB, text, tables, infoboxes).
CompassRank is dedicated to exploring the most advanced language and visual models, offering a comprehensive, objective, and neutral evaluation reference for the industry and research.
evaluates LLM's ability to call external functions/tools.
An Automatic Evaluator for Instruction-following Language Models using Nous benchmark suite.
aims to track, rank, and evaluate LLMs and chatbots as they are released.
a benchmark platform for large language models (LLMs) that features anonymous, randomized battles in a crowdsourced manner.
Pushing the frontier of cost-effective reasoning.
A collection of code samples from AWS which can be adapted for use with Claude. Note that some samples may require modification to work optimally with Claude.
## Contributing
, 2, 3
Kalman & Bayesian Filters in Python
Bayesian Reasoning and Deep Learning, Slides
Boosting vs Bagging
Ensembling models with R, Ensembling Regression Models in R, Intro to Ensembles in R