Very large online platforms (VLOPs) completed the required ‘harmonised’ reporting template for transparency in August, the second batch of reports since the template became mandatory in mid-2025, with the first harmonised reports landing in February 2026. These templates set a high standard for reporting, particularly for reporting by language. 

Are you prepared to report on all of your decisions by language? If you’re not already detecting and tracking language granularly with your moderation decisions, you may not be able to effectively produce the harmonised transparency template. 

What does this mean for VLOPs?

VLOPs are required to break down their reporting on automated actions by the language of the content that is actioned (you don't necessarily need to have 45 million registered accounts from the EU to qualify as a VLOP - users accessing without an account are included in the count for many online platforms). The EU’s instructions indicate that it should report based on the language of the content where the violation took place, even in nuanced multi-language content:

"For the calculation of measures imposed to specific items of information containing multiple languages, the language specified in the order or notice shall be predominant, or in case reference to a language is omitted, the predominant language of the infringing content shall be included. For example, the infringing item is a video in German with subtitles in English. If the infringing nature concerns the audio, then German is the predominant language. If the infringing nature concerns the text, English is the predominant language. If the infringement concerns both equally, the item is to be included in the calculations for both English and German. If the infringement does not concern the audio nor the text (e.g., solely the images), then it is not to be included in the calculation of language-specific accuracy".

For organisations operating with users speaking many European languages, delivering this reporting requires the moderation system to provide detailed tracking of the content’s language and the ability to associate any automated moderation decision with the part of content actioned and its correct language. Further, measures of recall, precision, and accuracy to judge the performance of automated actions should be broken down based on the language of the content itself too. A second challenge is tracking your teams of moderators to ensure you can report on the human resources dedicated to content moderation and their “qualifications and linguistic expertise” broken down by each official language. 

This is harder than it sounds! As described in the guidance quoted above, it’s not enough to simply report based on: 

  • the user’s location (e.g. to assume a German user’s content is in German)
  • the website (e.g. someone on a .fr site is speaking French)
  • part of the content language outside of the focus (e.g. a German video if it was an issue with an English subtitle)

It’s also easy to get into tough edge cases when trying to identify a language with granularity - particularly when messages are short. Consider the following chat messages:

  • “ ☀️hi!” - If this is in a French chat room, would you mark it as English?
  • “Schmuck” - Is this German for Jewelry or an insult in Yiddish?

The reporting requirement also doesn’t change depending on how many users you have in small-volume languages (e.g. your platform may have very little content in Maltese or Irish). To be compliant, these languages still need reporting and accuracy stats for automated action.

The result: companies are under-compliant. Platforms are struggling to provide this information, and even the largest and most established companies are inconsistent when it comes to language reporting. A study published in May 2026 showed that many platforms failed to comply with the harmonisation requirements on language. Of the eight companies studied, only 4 were compliant, and only two (Facebook and Instagram) covered all 24 official EU languages.

How can VLOP's meet the DSA language guidelines?

To meet these reporting guidelines, you need to manually or automatically label all components of your content with language.

 Doing this raises several key questions:

  1. How do you detect and label your content’s language at scale? 
    Across text, image, and video, you should track language metadata. With human moderators, they can manually tag these things (if they know the language), but at scale, this tagging can be done with a language detection service or a combination of language detection and user metadata to provide the most likely language for the content. If you’re a VLOP or suspect you may grow into one in the coming years, it is worth ensuring you are able to detect and report on language at scale. This can also become part of the metadata you provide within statements of reasons to the EU DSA Transparency database.
  2. How accurately can you detect language?

    Automatic language detection is ‘tricky’, not least because closely related languages can be difficult to distinguish (with research showing that precision collapses with very closely related languages). While they are imperfect tools, experimenting early with this process and understanding its limitations is key to delivering high-quality transparency reporting to meet your regulatory requirements.

  3. Are you prepared to audit and explain complex language stats?
    When an infringement concerns two languages equally, the item is counted twice (in both languages), but when content may have text and an image, if the image violates without language, it should not be counted against your language breakdown at all. This makes tracking these stats more critical to capture as part of your moderation system design instead of simply at reporting time.

While the tracking is painful, the principle behind it has good intentions. Reporting by language exists because moderation quality has never been evenly distributed across languages, and until now there was no way to prove it from the outside. Aggregate numbers hide language nuance completely (particularly for languages that represent small volumes). Researchers have already used the human-resource side of this reporting to show that some platforms had no moderators at all covering the national language of millions of their EU users, and that one major language received the same moderator count as another with a tiny fraction of the volume. As a trust and safety leader, the data also helps you understand which languages you may be weaker in before a regulator or a researcher tells you.

Complete the checklist below to see whether you are DSA-compliant for language moderation. 

 

Self-assessment

How ready are you for language detection and reporting?

0 / 11