Updated September 17, 2026 8:25 pm
In short
Baseten’s Base Labs is building an open-weight AI safety standard with Hugging Face and Goodfire, and is now inviting broader developer input as it tries to make safety evaluation and monitoring part of the open-model lifecycle.
- Baseten, Hugging Face and Goodfire AI are collaborating on open-weight model safety infrastructure.
- The partnership focuses on evaluation, monitoring and standards built into the model lifecycle.
- The move responds to the rise of abliteration, with more than 6,000 abliterated models listed on Hugging Face.
- Baseten says openness can improve safety by making model behavior more visible and more controllable.
- The company is inviting outside developers to help shape the framework.
Update — September 17, 2026 8:25 pm
Baseten also said it is opening the framework to outside developers, inviting the broader ecosystem to contribute to the effort as it takes shape.
The company additionally put fresh context around the partners, noting that Baseten raised a $1.5 billion Series F in June at a $13 billion valuation, while Goodfire recently secured a $150 million Series B to expand its interpretability work.
Baseten said on Wednesday that its Base Labs research unit is launching an open-weight AI safety partnership with Hugging Face and Goodfire AI, aiming to create evaluation and monitoring tools that make open models safer to train, deploy and serve. The move matters because the open-weight ecosystem now includes thousands of stripped-down models that can be made more dangerous when safety controls are removed.
The effort arrives as the AI industry wrestles with a growing problem: a technique known as “abliteration,” in which model safeguards are removed to turn a previously aligned system into something far less constrained. Hugging Face, one of the largest repositories for open-source models, now lists more than 6,000 abliterated models, underscoring how quickly safety weaknesses can spread once weights are available for public use.
Baseten is positioning the initiative as a broader framework rather than a one-off research project. The company says it wants open models to ship with safety practices that are visible, testable and integrated into the model lifecycle instead of being added later as an afterthought.
What Baseten announced and why it matters
Baseten’s announcement is really about infrastructure: the company wants to help define a common approach for safety in open-weight AI, with Base Labs publishing methods for training, evaluating and monitoring models. That is significant because open models have a major advantage over closed systems — anyone can inspect, adapt and deploy them — but that same openness can also make harmful modification easier.
Open-weight models are not inherently unsafe. In many cases, their transparency makes auditing easier, and researchers can test behavior more thoroughly than they can with proprietary systems. The concern is what happens after release, when bad actors or careless users remove safeguards and repurpose the model. Baseten’s pitch is that openness can be part of the solution, not just part of the problem.
In a post on X, the company said it believes openness is helpful for AI safety because it allows more visibility into model behavior and gives researchers more practical ways to turn safety findings into controls that are transparent and usable. The implication is that open models can be governed by technical measures that are more inspectable than the black-box safety layers typically used by closed providers.
How does the partnership work?
The partners have not published a detailed technical design, but their roles are easy to infer from their specialties. Baseten operates as an AI inference provider, which means it helps companies serve models efficiently at scale. Hugging Face is the open-model distribution hub where researchers and developers share and download models. Goodfire focuses on interpretability, the effort to understand how neural networks make decisions.
Together, the three companies could cover the full path from model release to deployment: Hugging Face as the distribution layer, Goodfire as the interpretability and monitoring layer, and Baseten as the serving and infrastructure layer. That combination suggests the group is trying to make safety measurable at each stage, not just during initial training.
Goodfire emphasized that safety should be designed into open models and maintained by the companies that host or serve them. That framing points to a broader policy question in AI: once a model is open-weight, who is responsible when it is modified, deployed or repackaged in a harmful way?
Baseten said openness creates more visibility into how models behave and more ways to translate safety research into practical controls than closed-source systems allow. Goodfire added that safety has to be built into open models and supported by the organizations that serve them.
Why open-weight model safety is becoming a bigger issue
Open-weight models are gaining influence because they let developers run powerful AI systems on their own infrastructure, customize them for specific tasks and avoid depending entirely on a single vendor. That flexibility has helped fuel a fast-growing ecosystem across startups, research groups and enterprise users.
But the same openness also means the model’s weights can be altered, redistributed and optimized for different purposes. Safety systems that are effective in a hosted product can disappear when someone downloads the underlying model and modifies it. That is the core issue behind the rise of abliteration, a practice that strips away guardrails and can transform an otherwise limited model into one that is much easier to misuse.
The scale of the problem is already substantial. Hugging Face currently shows more than 6,000 abliterated models, a sign that this is no longer a theoretical concern or a niche research edge case. It has become a broad ecosystem challenge, especially as open model development accelerates and as more organizations seek to host or fine-tune models independently.
That context helps explain why Baseten is framing safety as a standard rather than a feature. A standard can be adopted by multiple vendors, reused across deployments and updated as model capabilities evolve. A feature, by contrast, is often isolated inside one product and can be bypassed once the model leaves that environment.
What is abliteration?
Abliteration is a method for removing a model’s safety mechanisms so it behaves differently from the aligned version that was originally released. In practical terms, it can mean taking a model that was trained or tuned to avoid harmful responses and then modifying it so those restrictions no longer apply.
This matters because alignment work is not always durable once a model is open. If the weights are accessible, a motivated user can alter them, rediscover behavior patterns or create new variants that are harder to supervise. That gives the open-weight community a different risk profile from the one faced by closed frontier labs.
The growing number of abliterated models on Hugging Face suggests the practice has moved from obscure experimentation into a recognizable pattern. For safety researchers, that creates a need for tools that can detect, benchmark and monitor risky modifications before they spread widely.
Why interpretability is central
Interpretability is central because it helps answer the question of what a model is actually doing internally. If researchers can identify the circuits, features or activations associated with unsafe behavior, they can design better monitors and interventions. That is the niche Goodfire has been building around.
Goodfire’s work is especially relevant in open ecosystems, where the goal is not only to limit bad outputs but also to understand the mechanisms behind them. If the industry wants open models to be trustworthy at scale, it will need more than user-facing filters. It will need technical visibility into how and why a model behaves as it does.
Who are the companies behind the partnership?
The partnership brings together three companies with different but overlapping strengths in the AI stack. Baseten has become a major infrastructure player for running models in production. Hugging Face is the best-known open model platform. Goodfire is one of the newer names in model interpretability.
That mix is important because safety problems are often spread across the AI lifecycle. A model might be trained in one place, published somewhere else, fine-tuned by a third party and then served through a separate provider. Any serious safety framework for open-weight systems has to address all of those handoffs.
| Company | Main role in the partnership | Why it matters |
|---|---|---|
| Baseten / Base Labs | Research, model serving and safety framework development | Connects safety methods to real deployment infrastructure |
| Hugging Face | Open-model hosting and ecosystem distribution | Provides the public marketplace where models spread |
| Goodfire AI | Interpretability and internal model analysis | Helps explain and monitor model behavior |
Why is Baseten making this move now?
Baseten appears to be moving now because open-weight AI is entering a phase where deployment scale and safety concerns are rising at the same time. As more enterprises and developers adopt open models, the need for common safety tooling becomes more urgent.
There is also a business logic to the move. Baseten is not just a research company; it is an infrastructure provider that benefits when models are run in a predictable, secure and enterprise-friendly way. Helping establish a safety standard for open models could reinforce trust in the kinds of deployments Baseten wants to host.
The company’s timing also aligns with its broader investment in research. Base Labs was created earlier this year, and the new partnership suggests that the unit is intended to do more than publish papers. It is being used as a vehicle to shape best practices in the market.
In the background, Baseten’s own financial position gives it the capacity to pursue this kind of work. The company raised a $1.5 billion Series F in June, a massive round that pushed its valuation to $13 billion. That capital can support research initiatives that do not have an immediate revenue payoff but may strengthen the company’s strategic position over time.
How much money are the partners working with?
All three companies come into the partnership from strong financial positions, especially by startup standards. That matters because building safety infrastructure for open models is not a lightweight task; it requires research talent, compute resources, and time spent developing standards that the market may or may not adopt quickly.
Baseten’s June financing was one of the largest recent AI infrastructure raises. Goodfire, meanwhile, also secured substantial backing this year, reflecting investor interest in tools that make large models more understandable and controllable. The financial strength of both companies suggests the partnership is backed by enough runway to support long-term experimentation.
For Hugging Face, the value is more ecosystem-driven. As the leading destination for open models, it has a strong interest in preserving trust in the platform. If open-weight AI becomes associated with low-quality or unsafe model variants, that could damage adoption across the broader ecosystem.
| Date | Event | Why it matters |
|---|---|---|
| Earlier in 2026 | Baseten launches Base Labs | Creates a research arm focused on model methods and standards |
| Earlier in 2026 | Goodfire raises $150 million Series B | Funds interpretability work for model understanding and safety |
| June 2026 | Baseten raises $1.5 billion Series F | Strengthens Baseten’s infrastructure and research ambitions |
| September 17, 2026 | Safety partnership announced | Sets out a collaborative framework for open-weight model safety |
What could an open-model safety standard include?
A useful open-model safety standard would likely include tools for pre-release evaluation, post-release monitoring and deployment-time controls. That could mean benchmarking models for harmful behavior, checking for signs of unauthorized modification and tracking how models are used once they are served to end users.
It may also include better documentation for how models were trained and tuned, along with audit trails for changes made after release. For open-weight systems, provenance is crucial. If safety controls are to be trusted, users need to know where the model came from, what changed and who is responsible for serving it.
The challenge is that there is no single universal threat model. A model used in research, enterprise support or public chat all faces different risks. Any standard that hopes to matter must be flexible enough to address those differences while still being specific enough to be enforceable.
Potential building blocks
While the companies have not detailed the exact architecture, a practical safety framework could include:
- model behavior audits before publication
- red-team testing for jailbreaks and harmful responses
- monitoring for abliterated or tampered variants
- interpretability-based alerts for suspicious internal changes
- deployment controls that can be updated without retraining the full model
Those building blocks would not eliminate risk, but they could make open-weight systems easier to govern in practice.
Why openness and safety are no longer opposites
The old assumption in AI safety was that closed systems were easier to control because only the vendor could alter them. But that logic has been challenged by the rise of open-weight development, where transparency can actually improve auditing and accountability.
Baseten is arguing that openness can help safety precisely because more people can inspect the model, replicate findings and apply fixes. That is a compelling argument in a field where much of the most powerful AI remains hidden behind proprietary APIs. If the community can standardize safe open deployment, it may create a model that is both broadly accessible and easier to study.
Still, openness alone is not enough. The release of a model does not guarantee that downstream users will preserve its safeguards. Safety tools have to travel with the model, or at least remain attached to the infrastructure that serves it. That is the problem Baseten, Hugging Face and Goodfire are trying to solve.
Baseten is inviting outside developers to contribute to the framework, signaling that it wants the initiative to become ecosystem-wide rather than owned by one vendor.
What happens next?
The next step is likely a series of technical releases, collaborative guidelines or evaluation tools that make the partnership more concrete. Baseten has said it wants the broader developer ecosystem involved, which suggests the company is hoping other model builders, serving platforms and researchers will adopt or extend the framework.
If the effort gains traction, it could become part of the emerging standard stack for open-weight model deployment. If it does not, it may still influence how companies think about safety, especially among infrastructure providers that sit between model developers and end users.
Either way, the announcement reflects an important shift in AI governance. Safety is moving away from a purely theoretical conversation and toward an engineering discipline that has to work in public, across organizations and at the speed of model release cycles.
For open-weight AI, that shift may be unavoidable. The ecosystem is too large, too fast and too distributed for safety to depend on good intentions alone. Baseten’s partnership with Hugging Face and Goodfire is a sign that the industry is starting to treat safety as a shared technical layer — one that must be built, measured and maintained just like the models themselves.
Timeline of the key developments
Here is the sequence of events that led to the new partnership:
- Open-weight models continue expanding across research and commercial use.
- Abliteration spreads as a method for stripping safety controls from released models.
- Hugging Face reports thousands of abliterated models in its ecosystem.
- Baseten launches Base Labs to develop research-led model methods.
- Baseten, Hugging Face and Goodfire announce a joint effort to build open-model safety infrastructure.
The timeline shows a familiar pattern in AI: technical capability moves faster than the standards needed to govern it. This partnership is an attempt to narrow that gap before unsafe practices become even more deeply embedded in the open-model ecosystem.
Frequently asked questions
What did Baseten announce with Hugging Face and Goodfire?
Baseten announced an open-weight AI safety partnership through its Base Labs research arm. The collaboration is intended to build evaluation and monitoring infrastructure for open models, with the goal of making safety part of the training and deployment process rather than a later add-on.
Why is this partnership important for open-source AI models?
This partnership is important because open-weight models can be modified after release, including in ways that remove safeguards. By creating shared safety tools and standards, the companies are trying to reduce the risk that open models are repurposed or deployed in harmful forms.
What is abliteration in AI?
Abliteration is the practice of removing or weakening a model’s safety controls so it behaves differently from the aligned version that was originally released. In open-weight ecosystems, that can make it easier for users to create less restricted and potentially more dangerous variants.
How many abliterated models are on Hugging Face?
Hugging Face currently lists more than 6,000 abliterated models. That figure highlights how widespread the issue has become and why companies are pushing for better monitoring and safety standards for open-weight AI.
Who are Baseten, Hugging Face and Goodfire?
Baseten is an AI inference and infrastructure company, Hugging Face is a major hub for open-source models, and Goodfire specializes in interpretability research. Together, they cover model hosting, serving and internal analysis, which could make them well suited to an open-model safety framework.









