Startups

Hugging Face says breach exposed internal data and credentials

The AI platform says attackers used a malicious dataset to reach internal systems, and it is telling users to rotate stored access keys.

Theo Nakamura

By Theo Nakamura · Staff Writer

· 3 min read

Hugging Face says breach exposed internal data and credentials
Photo: TechCrunch

Hugging Face said a cyberattack reached its internal datasets and service credentials, a reminder that AI infrastructure now carries the same operational risk as any major cloud platform. For developers, startups and companies building with open models, the immediate takeaway is practical: check account activity and replace sensitive keys stored on the service.

The company disclosed the incident Friday in a blog post, according to TechCrunch. Hugging Face said it was still looking into whether customer or partner information was taken.

Hugging Face hosts AI models and datasets, which are the files and training materials developers use to build and run artificial intelligence systems. The company said the attack began with a dataset uploaded to its platform that exploited a security flaw and ran malicious code on Hugging Face servers.

That code let the attackers raise their level of access inside the company’s systems, Hugging Face said. In security terms, that is known as privilege escalation: an attacker starts with limited access, then uses a bug or misconfiguration to gain broader control.

What users are being asked to do

Hugging Face said it has fixed the vulnerability and revoked and replaced the credentials that were accessed. Credentials are login or service secrets that let software connect to systems. Access tokens, a common type of credential, work like digital keys for applications and developer accounts.

The company urged users to rotate any keys stored on Hugging Face and review accounts for unusual activity. Rotating a key means invalidating the old one and creating a new one, so anyone who copied the previous token can no longer use it.

Hugging Face said its own anomaly detection systems identified the attack. It also said it used an AI model to study server logs, which are records of activity on a company’s systems, to understand what happened.

According to the company, it first tried to use a frontier AI model from a commercial provider. A frontier model is a high-end, general-purpose AI system from a leading lab. Hugging Face said that effort was blocked by the provider’s safety guardrails, so it used its own local large language model instead. The company said that approach also avoided sending sensitive attack logs to another company’s servers.

Company says an outside AI agent was involved

Hugging Face attributed the breach to an external AI agent, saying it carried out “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.” A sandbox is an isolated computing environment used to run code with limits.

TechCrunch reported that Hugging Face did not immediately provide evidence for that claim when asked. The company has reported the incident to law enforcement and brought in cybersecurity forensic specialists to investigate the breach and review its defenses, according to TechCrunch.

The incident highlights a hard security problem for AI platforms: they invite users to upload and run code, models and datasets, while also needing to keep internal systems sealed off. Hugging Face said it has addressed the flaw used in this attack, but its investigation into possible data theft was still ongoing.

TechCrunch reported that it was unclear whether Hugging Face had completed a security audit before launch. A Hugging Face spokesperson did not respond to TechCrunch’s request for comment Monday.

This story draws on original reporting from TechCrunch.

More from Startups

All Startups