AI Robotics Ethics Society®

Aira-Instruct 🤗

Aira-Instruct, a new series of language models fine-tuned via instruction-tuning and RLHF, released in Portuguese and English.

Aira-Instruct 🤗

We just released an improved version of our language model, Aira. Aira has several iterations, ranging from closed-domain chatbots to open-domain chatbots fine-tuned via instruction-tuning and RLHF (Reinforcement Learning from Human Feedback).

This new version, Aira-Instruct, is a series of generative language models, ranging from 124M to 1.7B parameters, available in Portuguese and English.

We’re also releasing two reward models (used in RLHF): one built to evaluate the quality of our models’ generations (RewardModelPT), and another to help control the toxicity present in the model’s generations (ToxicityModelPT). Both models are available in Portuguese and English.

The datasets used to train all of the mentioned models, along with the training implementation, are also available on Hugging Face. 🤗

The Aira-Instruct series was developed to help researchers explore challenges related to the Alignment Problem. Since these are small-scale models (up to 1.7 billion parameters), they can be reproduced by individual researchers at a relatively low cost (~R$250.00).

Try our demo on the AIRES Playground or on Hugging Face!

The models and datasets developed here are part of the doctoral thesis of Nicholas Kluge, “Dynamic Normativity: Necessary and Sufficient Conditions for Outer Alignment.” This research is funded by CNPq, FAPERGS (Foundation for Research Support of the State of Rio Grande do Sul), DAAD (German Academic Exchange Service), PUCRS (Pontifical Catholic University of Rio Grande do Sul), and the University of Bonn.