Can AI Face Swap "Steal" Someone's Likeness? The Ethics of Face Data
The most important question isn't legal. It's where the training data came from in the first place.

Key takeaways
- Most discussion of AI and facial ethics focuses on a single output — whose face shows up in one generated image or video. That's a real question, but it's not the whole picture.A separate, earlier question matters just as much: whose faces trained the underlying models capable of recognizing and reconstructing faces in the first place, and whether those people ever knew or agreed to it.
- Several well-documented face-recognition datasets built over the past decade were assembled by scraping publicly posted photos from the web or social media, without individually notifying the people depicted — practices that later drew significant public criticism, lawsuits, and regulatory action.
- Once a model has been trained on a dataset like that, there's generally no practical way for an individual to confirm whether their own photo was included, or to have it meaningfully removed after the fact — which is a different, harder problem than the "does this one app train on my photos going forward" question we've covered elsewhere on this blog.
- The core ethical question isn't really about commercial harm — it's about consent and control. Using someone's likeness without their knowledge is a meaningful loss of control over their own image, independent of whether any money changed hands.
Most of what gets written about AI and facial ethics — including a few pieces on this blog — focuses on a single generated image or video: whose face appears in it, whether they agreed to it, what it's being used for. That's an important question, and we've covered it from a few angles already. But it sits on top of a deeper, less visible one: before any individual output exists, where did the underlying model's understanding of human faces actually come from, and did anyone ask permission for that?
Two different ethical questions, often conflated
It's worth separating these clearly, because they call for different answers:
- Output-level consent: did the specific person depicted in a specific generated image or video agree to that use? This is what we covered in pieces on deepfakes versus face swap apps and common legal scenarios with AI images.
- Training-data consent: did the people whose photos taught the underlying model how faces work in general ever know their images were being used that way? This is a much older, much less visible question, and it's the one this piece focuses on.
An app can score well on the first question — clear consent, obvious labeling, your own photo, your own choice — while the models underneath it were built years earlier on a dataset that scored poorly on the second. Both are real ethical questions, but they're not the same one, and conflating them tends to make the training-data problem invisible.

Where face-recognition training data has actually come from
This isn't a hypothetical concern. Several well-documented cases from the past decade illustrate how facial datasets were built, often without the knowledge of the people in them:
- One widely reported case involved a company that built a facial-recognition tool by scraping billions of publicly posted photos from social media and other websites, without asking the platforms or the individuals depicted, then sold access to search against that database. This drew substantial regulatory scrutiny, multiple lawsuits, and enforcement action in several jurisdictions.
- Academic and industry research datasets used to train and benchmark facial-recognition systems have, in a number of documented instances, been assembled by pulling large numbers of photos from platforms like Flickr or YouTube — chosen for having permissive-looking licenses at a technical level — without any individual notice to the people whose faces appeared in them. Some of these datasets were later withdrawn or restricted after public reporting drew attention to how they'd been compiled.
The common thread across these cases isn't a single company or a single bad actor — it's a pattern: large-scale face data was often collected because it was technically accessible, not because anyone involved had specifically agreed to have their face become part of a training set for a recognition or generation system.
Why this is harder to fix than "just delete the data"
This is where the training-data version of this problem gets genuinely difficult, in a way the "does this app train on my photos" question we covered in an earlier piece doesn't fully capture. When a modern app states that your photo is used only for one inference and not retained for training, that's a forward-looking, checkable commitment about what happens to a photo you're uploading today. The training-data ethics problem is backward-looking: it's about datasets assembled years ago, often before today's scrutiny existed, that already shaped how currently-deployed models understand faces in general.
Even when a dataset's general sourcing becomes public knowledge, there's usually no individual lookup tool that lets a specific person confirm whether their own photo was included, and even less way to meaningfully have a specific image's influence removed from a model that's already been trained on it. That's a structurally different, and in practice much harder, problem than checking a privacy policy before you upload a new photo today.
You can check whether an app trains on the photo you upload tomorrow. You generally can't check what already trained the model it's running on.
Why "was it commercial" isn't the right ethical test
Legal frameworks like the right of publicity, which we covered in an earlier piece on AI image legality, often center on commercial harm — did someone profit from using your likeness without permission. That's a reasonable place for the law to draw an enforceable line, but it's a narrower question than the ethical one. Using someone's face without their knowledge is a meaningful violation of their control over their own image whether or not any money ever changed hands. A dataset built from scraped social media photos, used purely for non-commercial research, still involved using people's faces without asking — the absence of a sale doesn't resolve the absence of consent.

This is really the same conceptual thread running through most of what we've written on this blog about faces and AI: consent, not commercial outcome, is the variable that actually determines whether a use is respectful of the person involved. It shows up in how biometric data gets treated differently from other personal data, in what separates a deepfake from a consumer app, and here, at the level of the training data itself.
What responsible practice actually looks like
None of this means facial-recognition or face-generation technology is inherently unethical — it means the ethics live in specific, checkable practices rather than in the technology as a category:
- Informed consent at the point of collection — people should generally know when their photo is being used to build or improve a system that recognizes or generates faces, rather than finding out after the fact through a news report.
- Transparency about data sourcing — companies disclosing, at least in general terms, where their training data comes from, rather than treating it as entirely proprietary and unexaminable.
- Minimizing unnecessary retention — the same "one inference, not retained for training" distinction we've covered before, applied going forward even if it can't undo past practices.
- Giving people real control — deletion rights, opt-outs, and honestly, in an ideal case, the ability to find out if you're in a given dataset at all, though this remains the least solved item on this list industry-wide.
Content provenance standards like the Content Credentials system we covered in an earlier piece on spotting AI-generated photos address a related but distinct problem — proving what happened to an image after it was created — rather than the training-data provenance question this piece is about, but both point toward the same broader shift: an industry slowly moving from "collect what's technically accessible" toward "be able to account for where this came from."
What this actually means if you're just using an app
As an individual user, the training-data ethics problem described here isn't something you can personally audit or fix — it's a systemic, industry-level question rather than a per-use decision. What is within your control is the narrower, output-level question we've covered elsewhere: whose face you choose to use, whether they've agreed to it, and whether the specific app you're using is transparent about what happens to the photo you upload today. Those two levels of the problem are connected, but they're not the same, and it's worth not letting the harder, systemic question make the more immediate, personal one feel unimportant by comparison.

A quick glossary
- Training data provenance
- Information about where a dataset used to train an AI model actually came from, and under what terms it was collected.
- Web scraping
- Automatically collecting publicly accessible content — including photos — from websites, often at large scale, without individually contacting the people or sites involved.
- Informed consent
- Agreement to a specific use of one's data or likeness, given with a genuine understanding of what that use involves — a higher bar than simply not objecting.
Frequently asked questions
Were AI face models trained on photos without people's permission?
In a number of well-documented cases, yes. Several face-recognition datasets built over the past decade were assembled by scraping publicly posted photos from the web or social media without individually notifying or asking the people depicted, which led to significant public controversy, lawsuits, and in some cases regulatory action.
Is it unethical to use an AI face swap app on my own photo?
Using your own photo, with your own consent, for your own entertainment isn't the ethical problem this topic is usually about. The concerns center on non-consensual use of someone else's likeness and on how the underlying training data for these systems was originally collected, not on personal use of your own face.
What's the difference between a legal violation and an ethical problem with face data?
Law sets an enforceable minimum, often centered on commercial harm or specific statutory categories. Ethics extends further, treating a person's control over their own image as valuable even when no money changes hands and no specific law is broken.
Can I find out if my photos were used to train a specific AI model?
Usually not with any certainty. Most large training datasets don't provide an individual lookup tool, and even when a dataset's general sourcing is publicly known, confirming whether one specific photo was included is often impractical after the fact.