Skip to content
Machine Learning Daily, home

AI Firms Tap Shuttered Businesses’ Data, Ignite Privacy Concerns

As public data sources dry up, AI models are now being trained on the internal communications of defunct companies, raising significant ethical questions about employee privacy.

DERRICKPOLICY & POWER722 WORDS
AI Firms Tap Shuttered Businesses’ Data, Ignite Privacy Concerns

The digital detritus of a defunct business, once relegated to the graveyard of discarded hard drives and forgotten cloud servers, has suddenly been resurrected as a valuable commodity in the ravenous maw of the artificial intelligence industry.

Shanna Johnson, CEO of the now-shuttered transcription and captioning company cielo24, discovered this unexpected new reality when winding down her 13-year-old enterprise.

What she termed the company’s “operational exhaust” – a sprawling archive of employee Slack conversations, Jira tickets, email correspondence, and multi-terabyte Google Drive files – became a surprising source of revenue, generating hundreds of thousands of dollars through a novel partnership with the wind-down specialist SimpleClosure.

This is not an isolated anecdote but a burgeoning trend, signaling a profound shift in how AI models are being trained and the increasingly intimate data sources they consume.

The public internet, once a seemingly boundless reservoir of information, has largely been drained.

Ilya Sutskever, a former chief scientist at OpenAI, reportedly suggested that by late 2024, AI companies would have consumed all readily available public content, including vast troves of Reddit discussions, Wikipedia entries, and digitized books.

This depletion, however, is only one part of the equation.

More critically, publicly available data, while extensive, often lacks the granularity and contextual “noise” necessary to forge truly effective, agentic AI systems – those designed to perform complex work tasks with human-like competence.

Enter the operational records of failed businesses.

These detailed digital footprints capture the authentic ebb and flow of workplace activities, decision-making processes, and the nuanced interactions that define daily work life.

Ali Ansari, whose company micro1 offers a product called Roots to AI laboratories, explains the imperative.

“Model companies are realizing the noise in the real-world environments is required to accurately test models,” Ansari noted.

His Roots platform, for instance, simulates a holding company where AI agents hone skills from financial services management to intricate calendar coordination, all based on the lived experience captured in corporate data.

This insatiable demand for authentic workplace data has turned SimpleClosure, originally a facilitator of standard company wind-down procedures like payroll termination and tax filings, into an unlikely data broker.

CEO Dori Yona describes the interest from AI companies as “insane,” likening the current climate to a “gold rush” for real-world operational data.

SimpleClosure is now poised to launch Asset Hub, a dedicated platform where closing companies can sell their code repositories, Slack archives, emails, and similar internal materials.

The company, which has completed nearly 100 such transactions over the past year, recovering over a million dollars for founders with payments typically ranging from $10,000 to $100,000 per company, acknowledges the sensitivity of this data.

Asset Hub is currently in beta testing, with SimpleClosure committing to remove all personally identifiable information (PII) – a technically challenging process that Yona insists must be perfected before widespread deployment.

Yet, despite assurances of PII removal, the practice ignites significant privacy concerns.

Marc Rotenberg, founder of the Center for AI and Digital Policy, has raised sharp questions about the ethical implications.

Employees, he argues, likely never anticipated that their internal communications, shared within the assumed confines of a workplace messaging system like Slack, would one day be repurposed as training fodder for AI.

While employees may sign intellectual property agreements covering work materials, the expectation of privacy in daily communication often transcends these legal clauses.

“I think the privacy issues here are quite substantial,” Rotenberg asserts, emphasizing that such data is not “generic” but involves “identifiable people.”

His organization has already alerted the Senate Commerce Committee, urging the Federal Trade Commission to scrutinize these novel AI business practices and reinforce personal data protection safeguards.

The ramifications extend far beyond defunct companies.

As newer and more powerful AI systems permeate every sector, the line between what is considered private and what is fair game for data harvesting becomes increasingly blurred.

Business communications, once implicitly private, are now viewed as crucial input for AI’s voracious appetite for raw data.

This raises profound questions about employee rights, the evolving nature of digital privacy in the workplace, and the regulatory frameworks needed to govern an industry racing to develop the next generation of intelligent agents.

Without robust protections and explicit transparency, the operational exhaust of today’s businesses could become the privacy nightmare of tomorrow’s workforce, inadvertently fueling AI’s ascent at the expense of individual autonomy.

REFERENCED

  1. cloud.google.comartificial intelligence industry
  2. simpleclosure.comSimpleClosure
  3. ibm.comagentic AI systems
  4. ibm.compersonally identifiable information (PII)

FILED TO POLICY & POWER

MORE IN POLICY & POWER

ALL

THE DISPATCH

Applied machine learning, filed daily.

Model releases, silicon, clinical deployment, and the policy shaping them. No digest padding.