Skip to main content
CIPP/CCIPP/CPrivacyCanadaPIPEDA

Joint Investigations and GenAI Scraping: Public Is Not Permission

Canadian joint investigations reject 'public web equals free training data.' Provenance, filtering, and contracts before GenAI scrape or fine-tune.

4 min read
ShareLinkedIn

Generative AI made an old privacy argument loud again: if something appears on the public internet, can a company vacuum it up for training?

Canadian regulators have been answering versions of that question for years. The short version is no—not without a valid legal framework, and not as a casual industry entitlement.

The precedent pattern: Clearview

The joint investigation into Clearview AI remains the clearest teaching case. Federal and provincial commissioners examined mass collection of images for a biometric system and rejected the idea that “publicly accessible” equals “up for industrial reuse.” The findings are public in the OPC’s PIPEDA Report of Findings 2021-001.

That file matters beyond facial recognition. It established cooperative enforcement muscle across jurisdictions and a substantive rejection of public-data exceptionalism for commercial AI.

Appellate reinforcement on extraterritorial reach in the Clearview litigation stream, including the B.C. Court of Appeal’s 2026 decision, adds judicial weight to the same ecosystem story: Canadian privacy law can follow scraping business models that affect people here.

Generative AI is the same conflict at larger scale

Large language models are trained on oceans of text and media that include personal information—names, contact details, biographical claims, children’s data, and sensitive inferences. Industry often responds with scale arguments: the corpus is too big to clean; the model does not “store” records; outputs are generative, not retrieval.

Regulators care about collection and use at the front end, accuracy and retention duties in the middle, and individual rights at the back end. “The dataset was big” is not a privacy principle under PIPEDA or stricter provincial statutes.

Through 2025–2026, multi-authority scrutiny of generative AI developers has followed the Clearview pattern: joint or coordinated pressure, focus on overbroad scraping, weak or nonexistent consent, and inadequate handling of sensitive and children’s information. Treat detailed company-specific findings as an evolving enforcement narrative grounded in that established approach, not as an excuse to invent tidy citation shortcuts. The compliance lesson does not require drama. It requires provenance.

What I require before model training or fine-tuning

Source inventory. Where did each corpus come from—licensed data, customer data, employee data, web scrape, vendor feed?

Personal information assessment. Did anyone actually check, or did the team assume code and blog posts only?

Lawful basis and notice. Especially for customer content and publicly posted personal details reused out of original context.

Filtering and minimization. Can you exclude private identifiers, credentials, and special-category data before training?

Retention and reuse rules. Is chat exhaust flowing back into weights without authority?

Accuracy controls. Hallucinated personal facts are not a cute demo bug under regimes that require accuracy of personal information.

Exit plan. If a regulator challenges a corpus, can you stop use, retrain, or retire a model path? Algorithmic sunk cost is not a legal defence.

Enterprise buyers are in the blast radius

Even if you do not train foundation models, you deploy them. You paste confidential data into prompts. You buy embeddings tools. You fine-tune on support logs.

Procurement questions I now treat as mandatory:

  • Training-data provenance summary
  • Opt-out and deletion realities for user content
  • Whether business data is used to improve shared models
  • Age and sensitive-data handling
  • Incident history with privacy regulators
  • Contract language on Canadian law and regulator cooperation

If a vendor cannot answer without marketing fog, price that as risk.

Connecting scraping risk to the wider Canadian agenda

This issue sits at the intersection of current PIPEDA accountability, provincial statutes such as Quebec’s Law 25, reform pressure after Bill C-27’s death, and global baselines like the EU AI Act. You do not need perfect harmony among those instruments to justify internal controls. You need humility about personal information in machine-learning supply chains.

I tell founders a blunt version: if your moat is “we scraped harder,” your moat is a liability. Licensed data, first-party relationships, and synthetic data with clear lineage are slower at the start and sturdier under investigation. The organizations that will struggle most are those that productized ambiguity—public posts, public images, public profiles—without ever asking whether commercial reuse matched the original context of publication.

Keep a living register of AI systems, corpora, and residual uncertainty. Uncertainty is allowed for a time. Undocumented uncertainty in production is how joint investigations turn into existential product problems.

A board-level narrative that works

Boards do not need a seminar on transformers. They need three sentences:

  1. Public web data can still be personal information when used commercially.
  2. Canadian authorities have already rejected “public equals free” for AI-scale collection.
  3. Our residual risk is whatever we cannot provenance, filter, or contractually control.

Then show the register. Executives fund what they can see. Hidden scrapes do not stay hidden once a coordinated investigation begins.

Actionable takeaway

Commission a 45-day training-data provenance review for every AI system that materially affects customers or employees. Classify each data source as licensed, first-party authorized, or scraped/unclear. Ban unclear sources from production training until legal and privacy sign off a written basis—or replace them. Anchor the review in the Clearview joint findings and your live duties under PIPEDA and, where relevant, Quebec’s CAI guidance. That single hygiene practice prevents more future pain than another abstract “responsible AI” statement.

Related services

Practical consulting aligned to this article’s focus—program design, controls, and operational delivery.

Browse all services