Training-data pipelines need secret scanning when public datasets can expose cloud, software, and AI-provider credentials at scale
Source: Truffle Security
Truffle Security scanned 7.6 petabytes of public Hugging Face datasets and reported 221,303 live, unique credentials across 6,003 datasets. The exposed material included software supply-chain tokens, cloud credentials, database logins, communications keys, and AI-provider keys. TLDR IT surfaced the research.
Why this matters: Treat data prepared for model training, evaluation, demos, and sharing as a credential-bearing supply-chain asset. Scan before publication and ingestion, use short-lived credentials with strict spending and scope limits, rotate exposed keys quickly, and keep provenance so a finding can be traced back to its owner and removal path.
Read the research