Skip to content
Machine Learning Daily, home

Cisco Releases Open-Source Model Provenance Kit as AI-BOMs Target Shadow Machine Learning

Security teams are pivoting from traditional software bills of materials to specialized frameworks that track model weights, agentic skills, and non-human identities.

DERRICKENTERPRISE & OPS806 WORDS

As enterprise architectures become increasingly saturated with unsanctioned machine learning models and agentic workflows, security teams are pivoting from traditional software bills of materials to specialized AI-BOMs to map hidden dependencies.

The traditional software bill of materials is no longer sufficient for modern MLOps pipelines, as it fails to capture the complex web of datasets, SDK libraries, and Model Context Protocol servers that define contemporary artificial intelligence applications.

While a standard SBOM captures static software packages, an AI-BOM must dynamically track ML frameworks, agentic skills, and the specific ways these elements interact within automated enterprise workflows.

To address this visibility gap, Cisco expanded its open-source AI security portfolio on Friday with the release of the Model Provenance Kit, a diagnostic tool designed to verify the lineage of deployed machine learning assets.

The repository operates through two primary mechanisms: a compare mode that evaluates metadata, tokenizer structures, and weight-level signals between two models, and a scan mode that matches a single model against a comprehensive fingerprint database.

This fingerprint repository currently encompasses approximately 150 base models across more than 45 architectural families and 20 publishers, providing a verifiable method to trace fine-tuned iterations back to their foundation models.

Amy Chang, head of AI threat intelligence and security research at Cisco, explained to The Register that the toolkit performs critical gate checks to delineate provenance-linked relationships.

“First, at the metadata level, it compares the information from the base model with the fine-tuned version of the model to delineate some sort of provenance-linked relationship – like this was derived from Meta Llama 4, or derived from Alibaba Qwen3,” Chang said.

By analyzing weight-based signifiers, the open-source tool provides a repeatable method to attest that customer-facing models operating within an enterprise environment align with the organization’s established risk tolerance.

Beyond model weights, comprehensive AI-BOM frameworks must now inventory the specific prompts driving agentic behaviors, as well as the underlying infrastructure utilized during the training phase.

Ziad Ghalleb, technical product marketing manager at Google’s Wiz, noted that effective AI-BOMs must also account for the non-human identities and permission sets attached to these automated workloads.

Because modern AI architectures rely heavily on autonomous agents interacting with various enterprise APIs, tracking the specific execution privileges of these non-human identities is just as critical as scanning the model parameters.

This expanded visibility extends to the developer workstations and integrated development environments where the AI applications are initially compiled, ensuring that every tool in the pipeline is logically tracked.

The push for granular ML observability stems from the rapid proliferation of shadow AI, where developers integrate unsanctioned coding assistants and external APIs without centralized security oversight.

Ian Swanson, vice president of AI security at Palo Alto Networks, characterized this visibility gap as a critical vulnerability, noting that organizations cannot protect architectures when they lack an understanding of the underlying components and their systemic interactions.

Imagine if AI is a birthday cake in the middle of this room, but you don’t know how it got there. You don’t know the recipe, you don’t know the ingredients, you don’t know the baker. Would you eat a slice of that cake?

Unvetted model integration carries immediate regulatory and compliance risks, particularly as frameworks like the European Union’s AI Act mandate strict documentation of training methodologies and data provenance for high-risk systems.

Chang highlighted recent industry examples, such as the integration of the Chinese open-source model Kimi 2.5 into Cursor’s Composer 2, as instances where obscured model lineage could trigger compliance violations or data sovereignty concerns.

The lack of state-level tracking also exposes enterprises to sophisticated adversarial tactics, including automated reconnaissance and prompt manipulation executed by malicious actors.

Criminal syndicates are actively leveraging these same machine learning frameworks to accelerate their operational tempo, utilizing automated agents to map compromised infrastructure and identify exposed endpoints.

Swanson detailed a recent incident investigated by Palo Alto Networks where attackers compromised an organization’s internal AI system prompts, modifying the instructions to force the model to exfiltrate sensitive data to an external environment.

If security teams possess a baseline understanding of the AI system’s state through an AI-BOM, they can rapidly identify when a core system prompt has been maliciously altered from its original configuration.

An operational AI-BOM provides a point-in-time snapshot of system configurations, allowing security teams to detect unauthorized state changes in system prompts or agentic skills before data exfiltration occurs.

As threat actors increasingly deploy automated agents to scan network blocks and manage attack infrastructure, the defensive requirement for real-time model telemetry will only intensify.

Sherrod DeGrippo, general manager of global threat intelligence at Microsoft, warned that agentic reconnaissance against enterprise systems represents a growing vector that requires immediate defensive mapping.

Future iterations of AI-BOM standards will likely need to integrate dynamic behavioral monitoring to detect runtime skills poisoning alongside the static component analysis currently being deployed.

REFERENCED

  1. cisa.govsoftware bill of materials
  2. securityweek.comModel Provenance Kit
  3. go.theregister.comThe Register
  4. ai.meta.comMeta Llama 4
  5. alibabacloud.comAlibaba Qwen3
  6. eur-lex.europa.euEuropean Union’s AI Act
  7. microsoft.comMicrosoft

FILED TO ENTERPRISE & OPS

MORE IN ENTERPRISE & OPS

ALL

THE DISPATCH

Applied machine learning, filed daily.

Model releases, silicon, clinical deployment, and the policy shaping them. No digest padding.