Baseten launches open AI model safety partnership — TechCrunch
Baseten, together with its research division Base Labs, announced a partnership with Hugging Face and Goodfire AI to create infrastructure for evaluating and monitoring the safety of AI models with open weights. As TechCrunch reports, Base Labs will develop and publish methods for training and monitoring such models.
A standard for open models
Baseten positions the future work as a safety standard for open models. According to the company’s plan, control mechanisms should be transparent and integrated into the training and deployment of models rather than added later.
The company said that openness is an advantage for artificial intelligence safety because it provides more opportunities to observe model behavior and turn safety research into practical and transparent control mechanisms.
More current news is available on the UA.News Telegram channel Telegram.
The problem of removing restrictions
The announcement comes amid discussions about the safety of models with open weights. The article refers to the abliteration technique, which can be used to remove a model’s built-in safeguards, potentially making it dangerous. Hugging Face, which hosts open-source AI models, currently has more than 6,000 models to which abliteration has been applied.
The companies did not disclose technical details of the collaboration. Goodfire AI, which specializes in model interpretability and explaining their decision-making mechanisms, said in response to Baseten’s post that safety must be built into open models and ensured by those who operate them.
Baseten also called on developers to join in shaping this approach. The company said it seeks to create an ecosystem of open models that are safe and accessible to everyone.