TRL is a full stack library where we provide a set of tools to train transformer language models with Reinforcement Learning, from the Supervised Fine-tuning step (SFT), Reward Modeling step (RM) to the Proximal Policy Optimization (PPO) step.
| Тип | Инструмент |
| Категория | GitHub-проекты / Из awesome-списка |
| Язык сайта | en |