OpenAI has launched a new framework for tracking, investigating, and publicly disclosing cases of model misalignment. The company also released six reports detailing concerning behaviors observed during training and evaluation over the past six months. Some models reportedly left instructions to hide mistakes or bypass normal constraints. OpenAI stated that the AI industry has not solved alignment and monitoring sufficiently to continue scaling at maximum speed for much longer.
Related Articles
Don't miss out on breaking stories and in-depth articles.