In section Releases

Anthropic Endorses Modular Training to Secure Open-Weights AI

Anthropic CEO Dario Amodei has identified joint research from AE Studio and Anthropic as a pivotal step toward securing open-weights AI models, offering a technical solution to the long-standing dilemma between public accessibility and the risk of releasing inherently dangerous capabilities.

Anthropic Endorses Modular Training to Secure Open-Weights AI

The collaboration centers on modular training, a strategy that isolates sensitive data—such as advanced cyberattack techniques or virology—into discrete segments. By sequestering this information, developers can strip away high-risk modules before a model is released to the public, ensuring that the remaining system retains its functional utility without carrying the weight of hazardous knowledge. Testing indicates that models trained through this modular approach match the performance of those trained conventionally.

Judd Rosenblatt, CEO of AE Studio, noted that the methodology, known as GRAM, prevents jailbreaking because the underlying model never encodes the dangerous data in the first place. The research, led by Ethan Roland, Murat Cubuktepe, and Erick Martinez, enters the discourse at a time when policymakers are grappling with the governance of open-weights systems. Amodei emphasized that safety risks should be empirically determined through testing rather than preemptive bans, positioning modular architecture as a viable middle ground for the future of AI development.

Share:on TelegramXFacebook

Subscribe to our newsletter

Once a week — the best stories from our editors, no ads or push notifications. Delivered Sunday morning.

Comments (0)

Leave a comment

No comments yet. Be the first!