a benchmark platform designed for evaluating large language models (LLMs) on a range of tasks, particularly focusing on their performance in different aspects such as natural language understanding, reasoning, and generalization.
| Тип | Инструмент |
| Категория | GitHub-проекты / Из awesome-списка |