An estimated 1.95 trillion photos (more than 5 billion a day) will be taken in 2026, according to Rise Above Research, making it harder than ever for human moderators to review all the harmful or toxic content that comes out. And volume is only half of it: some of that content is graphic, abusive, or illegal.
No human team can get through millions or billions of images a day, and a 2022 study of content moderators linked their work with compassion fatigue. AI image moderation is the automated process of scanning, analyzing and filtering images uploaded to platforms, and it helps you find images that break your policies. No automated system catches every one.
Done well, AI image moderation can catch harmful images sooner and flag fewer safe ones by mistake. That can also cut handling time and ease the load on human moderators.
AI Image moderation at a glance
- AI image moderation is the automated scanning, analysis and filtering of images uploaded to a platform to find content that breaks its policies, such as nudity, violence, hate symbols or drugs.
- Human moderators cannot review the volume of images platforms receive, and researchers have linked moderation work with compassion fatigue.
- Image moderation is hard because an image’s meaning often depends on context, such as captions or surrounding text, and bad actors manipulate images to evade detection.
- AI image moderation works by analyzing each image, labeling what it shows against the platform’s policies, scoring how likely a violation is and acting on that score.
- Confidence thresholds set the cutoff points that decide whether an image is actioned automatically, sent to a human moderator for review or allowed through.
- A common approach to image moderation combines AI and human moderators: AI handles the volume, and humans bring the context and judgment needed for borderline cases and appeals.
- The EU’s Digital Services Act and the UK’s Online Safety Act both cover illegal image content. The DSA focuses on transparency and user reporting, and the OSA adds proactive duties and age checks for services likely to be accessed by children.
What makes image moderation hard to get right
What makes image moderation hard is all the context an image can leave out. A picture may paint a thousand words, but context can flip that story on its head.
One very graphic image could be raising awareness for a medical fundraiser, and an innocent-looking image can carry a hateful message once text is added. Captions and overlaid or surrounding text heavily influence moderation decisions: If they don’t match the image, it can cause confusion.
Lack of context and gaps in how models are trained can lead to multiple issues:
- Malicious users bypassing moderation with manipulated images
- Safe images being incorrectly blocked (false positives)
- Models trained on biased datasets
- Edited and AI-generated images being misread, such as overlays, memes, obfuscated images and cropped or rotated copies
What works well is a mix of AI models and human moderators. Let the AI do the heavy lifting, which is sifting through potentially billions of images. Then, when the AI is unsure or an appeal comes through, a human moderator can do quality assurance and decide to block something or let it back onto the platform.
Doing it this way cuts the number of images people must review and lets moderators focus on the cases that need judgment.
What is AI image moderation and how does it actually work?
AI image moderation is built on models trained on datasets that define both acceptable and unacceptable imagery. By learning patterns, shapes, colors and contextual cues, these models can decide if content contains nudity, violence, hate symbols, drugs or other policy violations.
The 4 steps of AI image moderation typically look like this:
- Image upload: When a user uploads an image, it is sent through an API into the AI moderation system.
- Image labeling and categorization: Computer vision classifiers analyze the image and assign labels from a moderation taxonomy, and some tools can also locate individual objects, such as a weapon.
- Scoring and classification: The system gives each image a confidence score for how likely it is to break each policy, such as a percentage from 0 to 100 in Amazon Rekognition. Some tools use severity levels instead.
- Action: The AI can decide how to deal with the image. Ambiguous cases or appeals are sent to human moderators for review.
Detecting and classifying content
While there are lots of different types of moderation, successful moderation always requires one thing: a detailed policy for AI and human moderators to categorize content with. This determines what content is acceptable within your platform. Only once you have determined these policies can you come up with a logical system for detecting and assessing content.
Depending on the tool, AI image moderation can identify a wide range of potentially harmful or unwanted content, such as:
- Nudity
- Sexual or racy content
- Violence and gore
- Weapons
- Drugs
- Alcohol
- Hate symbols and hateful imagery
- Offensive gestures
- Gambling
- Profane or offensive text
- Faces (including estimated age range and apparent gender)
- Celebrity likenesses
- QR codes
- Images already published elsewhere online
Confidence scoring and thresholds
A confidence score is the model’s estimate of how likely an image is to break a specific policy. By default, Amazon Rekognition flags a label only when it’s at least 50% confident, according to its moderation documentation. A common setup gives each policy rule two thresholds: a higher one for automatic action and a lower one for human review.
Routing works in four steps:
- Score assignment: AI systems assign confidence scores to each content prediction.
- Threshold comparison: Each confidence score is compared against the policy’s two thresholds.
- Routing decision: Content at or above the higher threshold goes to automated action, and content below the lower threshold isn’t flagged. Content in between is routed for manual review.
- Processing execution: Content is either actioned automatically or flagged to human moderators.
Lowering the threshold for automatic action takes more harmful images down without review, along with more legitimate images that no person checks. Lowering the threshold for human review sends more borderline images to moderators, so fewer harmful ones go unflagged, but the queue grows.
What AI image moderation looks like in practice
JustGiving shows how AI and human moderators split the work on a platform where sensitive content is part of the purpose. The FCA-regulated fundraising platform hosts content some visitors will find upsetting, because it raises awareness of the causes people are fundraising for.
AI handles the volume. Checkstep flags content at scale, and JustGiving also uses those flags to spot potentially vulnerable customers.
Human moderators handle the context, because fundraising pages often use highly emotional language that a model would flag as a violation in other spaces.
JustGiving treats every false positive as something to learn from. Amy Somerford, Trust and Safety Lead at JustGiving, told us in a recent interview that the team keeps fine-tuning its AI and its rules to keep pace with new trends. As she put it: “The more we understand what’s on the platform, not just what the AI model is learning - but what we’re enforcing and not enforcing - the better, and that’s true for our moderators too.”
How Checkstep approaches AI image moderation
On Checkstep’s AI content moderation platform, “How good is my AI?” is a question our users can answer, because our model marketplace gives them real-time feedback on each model they run. That visibility is what makes customization possible, and customization is the key to good moderation.
Each part of Checkstep’s platform does a different job:
- A marketplace of models: Rather than running every image through one fixed, proprietary model, Checkstep gives platforms access to a marketplace of models. Different models suit different content types.
- Customer control: Customers control policies, thresholds and which AI model types touch their content, so they can optimize moderation for their specific use case. Tuning to the platform’s own content can reduce false positives.
- Advanced ModBot for ambiguous cases: Advanced ModBot, Checkstep’s AI reasoning model, makes decisions on suspicious content and returns its reasoning. It escalates to a human moderator when it’s not sure, and can reduce how much content people need to review.
- Human review: When the AI flags content as ambiguous, or a customer’s workflow calls for it, that content goes to a human moderator for review. Context is often what separates borderline content from genuinely harmful content, and a model working alone can miss that. Keeping the human-in-the-loop layer by routing ambiguous cases to a person means harmful content is less likely to slip through.
- Audit trail and security: Checkstep keeps a full, DSA-ready audit trail of moderation decisions. Underpinning this is the same security posture that applies across the platform: SOC 2 Type II certification, routine third-party security reviews, continuous control monitoring through Drata and data encrypted in transit and at rest.
What the DSA and OSA require for image content
The EU’s Digital Services Act and the UK’s Online Safety Act both cover how platforms handle illegal image content, but they take different approaches. The DSA leans on transparency and notice-and-action, while the OSA adds proactive duties and, for services likely to be accessed by children, highly effective age checks.
What the DSA requires
- Notice and action (Article 16): systems that let anyone flag content they consider illegal, such as child sexual abuse material (CSAM), non-consensual intimate imagery, IP infringement or illegal hate speech.
- Statement of reasons (Article 17): an explanation to the affected user when content is removed, disabled, demoted or otherwise restricted, or an account is suspended.
- Reporting suspected crimes (Article 18): a duty to promptly alert law enforcement to a suspected criminal offense threatening someone’s life or safety, such as child sexual abuse.
- Extra duties for very large platforms (Articles 33, 34, 35, 37 and 39): platforms the European Commission designates for having 45 million or more average monthly active EU users face systemic risk assessments (Article 34) and mitigation measures, which can include prominently marking generated or manipulated images, such as deepfakes (Article 35(1)(k)). They also face independent audits (Article 37) and ad-repository rules when images appear in ads (Article 39).
What the OSA requires
- Risk assessments (sections 9 and 11): platforms must assess the risk of illegal image content and, for services likely to be accessed by children, harmful image content too.
- Age assurance (sections 12 and 81): highly effective age checks that keep children from pornography, self-harm material and other primary priority content, on services likely to be accessed by children and on pornography services.
- Layered detection: Layered detection: Ofcom’s codes push larger and file-sharing services at risk of image-based CSAM toward a layered setup: hash matching against known CSAM, human review of a proportion of matches and user reports. Ofcom has also proposed that some services use automated tools, where effective, to catch new CSAM that hasn’t been hashed yet, and the UK government’s June 2026 explanatory memorandum says Ofcom plans its final decision in fall 2026.
- Technology notices (section 121): where Ofcom judges it necessary and proportionate, it can require a platform to use accredited technology to remove publicly shared terrorism content, or child sexual exploitation and abuse (CSEA) content shared publicly or privately.
For operators, both laws mean being able to explain moderation decisions and keep records. The DSA requires statements of reasons for most restrictions (Article 17), and the OSA requires written records of risk assessments and compliance measures (section 23). Checkstep supports compliance with DSA-ready audit infrastructure and compliance reporting. For the full breakdown, see our DSA guide.
Tool or managed service: Choosing your image moderation model
The right image moderation model depends on who will run it. A tool or API suits teams with in-house Trust & Safety staff and engineering capacity. A managed service suits platforms that need the work done for them. Many platforms land between the two.
- A moderation tool or API fits when you have your own moderators, engineers to integrate and tune models, and clear policies to configure them against.
- A managed service fits when you don't have in-house reviewers, need round-the-clock coverage, or want experienced moderators running the queue.
- A platform approach combines the two, with models, policies, thresholds, human review workflows, and audit trails in one place, and your own team or a partner handling review.
Whichever model you choose, most image moderation setups combine several tools:
- General image detection: Cloud provider APIs such as Amazon Rekognition are one widely documented option, and Rekognition uses the same label taxonomy for images and stored video.
- An added layer of coverage: Some teams also run content through a large language model (LLM), such as OpenAI’s GPT models or Anthropic’s Claude, both of which can analyze images.
- Known CSAM: Shield by Project Arachnid, from the Canadian Centre for Child Protection, is a no-cost API that lets electronic service providers check images against known child sexual abuse material.
- Known and unknown CSAM: Athena by Resolver combines hash matching for known CSAM with an AI classifier that detects new and AI-generated material that has not yet been hashed.
Beyond these core tools, many other models are available. The right combination depends on your risk tolerance, budget and the kind of content you moderate.
Frequently asked questions about image moderation
What does it mean when an image is moderated?
When an image is moderated, it means someone or something checked it against the platform's rules to see if it breaks guidelines around things like nudity, violence, or hate symbols. If it does, the platform can remove it, blur it, or stop it from being shown or posted.
What's an example of image moderation in practice?
An example of image moderation in practice is a fundraising photo of a surgical wound. The AI scores it high for graphic content, but not high enough to remove automatically, so it goes to a human moderator. The moderator sees the fundraising context and approves it, keeping a legitimate appeal live.
Is content moderation a good thing or does it go too far?
Content moderation is necessary, but it doesn’t always get it right. Context gets missed, AI systems get manipulated or policies aren’t clear enough, so some harmful content slips through. It can also go too far and remove legitimate posts, which is why appeals matter and why the DSA requires platforms to give users a reason for most removals.
What does moderation mean?
Moderation, in the context Checkstep operates in, means reviewing and making decisions about user-generated or platform content. Moderators decide what stays up, gets removed, gets labeled or gets escalated, to keep online spaces safe and compliant while preserving legitimate speech.
Where to start with AI image moderation
With AI image moderation, start by considering what resources you already have and how much you need to add. Every platform that hosts user images for a significant number of EU or UK users has duties under the DSA or OSA, whatever its size, though the DSA spares micro and small businesses some obligations. A platform with thousands of daily uploads will need automation to keep up.
A strategic, healthy moderation system isn’t something you can build overnight. Building your own tools takes sustained investment in labeling data and tuning models against your policies. The Trust & Safety industry is full of expert tools and communities that have tested moderation across a huge variety of situations and contexts, and working with them can help you avoid many of the early mistakes.
Four practical first steps:
- Write or review your image policy, so you know exactly what you are detecting.
- Estimate your volume and risk, including whether children use your platform.
- Decide who will review flagged images: your own team, a partner, or both.
- Check what regulators expect you to record and report, such as statements of reasons under the DSA.
Checkstep combines multi-model AI detection, human review workflows and DSA-ready audit trails, so your team can moderate images at scale without losing the context that matters. If you want to see how this could work on your platform, talk to our team.