a benchmark that evaluates large language models on a variety of multimodal reasoning tasks, including language, natural and social sciences, physical and social commonsense, temporal reasoning, algebra, and geometry.
| Тип | Инструмент |
| Категория | GitHub-проекты / Из awesome-списка |
| Язык сайта | en |