Movcl

· By Marcus Reyes

AI Explained

Can AI Face Swap "Steal" Someone's Likeness? The Ethics of Face Data

The most important question isn't legal. It's where the training data came from in the first place.

Key takeaways

  • Most discussion of AI and facial ethics focuses on a single output — whose face shows up in one generated image or video. That's a real question, but it's not the whole picture.A separate, earlier question matters just as much: whose faces trained the underlying models capable of recognizing and reconstructing faces in the first place, and whether those people ever knew or agreed to it.
  • Several well-documented face-recognition datasets built over the past decade were assembled by scraping publicly posted photos from the web or social media, without individually notifying the people depicted — practices that later drew significant public criticism, lawsuits, and regulatory action.
  • Once a model has been trained on a dataset like that, there's generally no practical way for an individual to confirm whether their own photo was included, or to have it meaningfully removed after the fact — which is a different, harder problem than the "does this one app train on my photos going forward" question we've covered elsewhere on this blog.
  • The core ethical question isn't really about commercial harm — it's about consent and control. Using someone's likeness without their knowledge is a meaningful loss of control over their own image, independent of whether any money changed hands.

Most of what gets written about AI and facial ethics — including a few pieces on this blog — focuses on a single generated image or video: whose face appears in it, whether they agreed to it, what it's being used for. That's an important question, and we've covered it from a few angles already. But it sits on top of a deeper, less visible one: before any individual output exists, where did the underlying model's understanding of human faces actually come from, and did anyone ask permission for that?

Two different ethical questions, often conflated

It's worth separating these clearly, because they call for different answers:

An app can score well on the first question — clear consent, obvious labeling, your own photo, your own choice — while the models underneath it were built years earlier on a dataset that scored poorly on the second. Both are real ethical questions, but they're not the same one, and conflating them tends to make the training-data problem invisible.

Where face-recognition training data has actually come from

This isn't a hypothetical concern. Several well-documented cases from the past decade illustrate how facial datasets were built, often without the knowledge of the people in them:

The common thread across these cases isn't a single company or a single bad actor — it's a pattern: large-scale face data was often collected because it was technically accessible, not because anyone involved had specifically agreed to have their face become part of a training set for a recognition or generation system.

Why this is harder to fix than "just delete the data"

This is where the training-data version of this problem gets genuinely difficult, in a way the "does this app train on my photos" question we covered in an earlier piece doesn't fully capture. When a modern app states that your photo is used only for one inference and not retained for training, that's a forward-looking, checkable commitment about what happens to a photo you're uploading today. The training-data ethics problem is backward-looking: it's about datasets assembled years ago, often before today's scrutiny existed, that already shaped how currently-deployed models understand faces in general.

Even when a dataset's general sourcing becomes public knowledge, there's usually no individual lookup tool that lets a specific person confirm whether their own photo was included, and even less way to meaningfully have a specific image's influence removed from a model that's already been trained on it. That's a structurally different, and in practice much harder, problem than checking a privacy policy before you upload a new photo today.

You can check whether an app trains on the photo you upload tomorrow. You generally can't check what already trained the model it's running on.

Why "was it commercial" isn't the right ethical test

Legal frameworks like the right of publicity, which we covered in an earlier piece on AI image legality, often center on commercial harm — did someone profit from using your likeness without permission. That's a reasonable place for the law to draw an enforceable line, but it's a narrower question than the ethical one. Using someone's face without their knowledge is a meaningful violation of their control over their own image whether or not any money ever changed hands. A dataset built from scraped social media photos, used purely for non-commercial research, still involved using people's faces without asking — the absence of a sale doesn't resolve the absence of consent.

This is really the same conceptual thread running through most of what we've written on this blog about faces and AI: consent, not commercial outcome, is the variable that actually determines whether a use is respectful of the person involved. It shows up in how biometric data gets treated differently from other personal data, in what separates a deepfake from a consumer app, and here, at the level of the training data itself.

What responsible practice actually looks like

None of this means facial-recognition or face-generation technology is inherently unethical — it means the ethics live in specific, checkable practices rather than in the technology as a category:

Content provenance standards like the Content Credentials system we covered in an earlier piece on spotting AI-generated photos address a related but distinct problem — proving what happened to an image after it was created — rather than the training-data provenance question this piece is about, but both point toward the same broader shift: an industry slowly moving from "collect what's technically accessible" toward "be able to account for where this came from."

What this actually means if you're just using an app

As an individual user, the training-data ethics problem described here isn't something you can personally audit or fix — it's a systemic, industry-level question rather than a per-use decision. What is within your control is the narrower, output-level question we've covered elsewhere: whose face you choose to use, whether they've agreed to it, and whether the specific app you're using is transparent about what happens to the photo you upload today. Those two levels of the problem are connected, but they're not the same, and it's worth not letting the harder, systemic question make the more immediate, personal one feel unimportant by comparison.

A quick glossary

Training data provenance
Information about where a dataset used to train an AI model actually came from, and under what terms it was collected.
Web scraping
Automatically collecting publicly accessible content — including photos — from websites, often at large scale, without individually contacting the people or sites involved.
Informed consent
Agreement to a specific use of one's data or likeness, given with a genuine understanding of what that use involves — a higher bar than simply not objecting.

Frequently asked questions

Were AI face models trained on photos without people's permission?

In a number of well-documented cases, yes. Several face-recognition datasets built over the past decade were assembled by scraping publicly posted photos from the web or social media without individually notifying or asking the people depicted, which led to significant public controversy, lawsuits, and in some cases regulatory action.

Is it unethical to use an AI face swap app on my own photo?

Using your own photo, with your own consent, for your own entertainment isn't the ethical problem this topic is usually about. The concerns center on non-consensual use of someone else's likeness and on how the underlying training data for these systems was originally collected, not on personal use of your own face.

What's the difference between a legal violation and an ethical problem with face data?

Law sets an enforceable minimum, often centered on commercial harm or specific statutory categories. Ethics extends further, treating a person's control over their own image as valuable even when no money changes hands and no specific law is broken.

Can I find out if my photos were used to train a specific AI model?

Usually not with any certainty. Most large training datasets don't provide an individual lookup tool, and even when a dataset's general sourcing is publicly known, confirming whether one specific photo was included is often impractical after the fact.

A note on this piece: The dataset and company examples referenced here reflect widely reported, publicly documented controversies rather than a live, independently verified investigation on our part — specific details, outcomes, and legal statuses of these cases can change or have continued to develop, so verify current specifics against current reporting rather than treating this as a definitive account. This was written without access to real-time search.

About Marcus Reyes

Marcus writes about how generative AI actually works, in plain English. Former machine learning engineer, now translating research into things regular people can understand.