AI Act Training Summaries Fail to List Scraped Websites, Review Finds
A review of 24 publicly available AI Act training‑content summaries shows that none of the documents list scraped domains in the dedicated section, raising questions about compliance with the European Commission’s template.

ActuIA conducted a systematic review of twenty‑four AI Act training‑content summaries that were publicly accessible as of September 25, 2026. Each of these documents was produced in compliance with the reporting obligations set out in Article 53(1)(d) of the European Union’s AI Act, which mandates a detailed accounting of data‑scraping activities.
According to the European Commission’s prescribed template, providers must disclose the first‑level and second‑level domain names that together constitute the top ten percent of all scraped websites by volume. The threshold is lowered for small and medium‑sized enterprises, meaning that even modest players are required to list the most frequently harvested domains.
Majority of Summaries Omit Domain Lists
In the sample examined, twenty‑one of the twenty‑four summaries featured the dedicated section for major scraped domains, yet the field was left entirely blank. No website names were supplied, effectively rendering the required disclosure empty.
A distinct pattern emerged for DeepSeek, whose three submitted summaries omitted the web‑scraping section altogether. By excluding the section, the providers failed to provide any of the information that the template explicitly calls for.
How Providers Describe Their Data Sources
Rather than enumerating concrete domain names, the majority of the documents resorted to broad, generic categories. Typical phrasing included references to "academic resources," "code‑hosting platforms," "encyclopedias," "regional portals" and generic top‑level domains such as ".com" or ".org." This approach sidesteps the precise identification of individual sites.
The set of providers covered by the review spans a diverse cross‑section of the AI landscape. It includes well‑known entities such as OpenAI, Mistral AI, Z.AI, ByteDance, MiniMax, the three DeepSeek submissions, and a batch of eleven new documents submitted by Ant Group.
Ant Group’s Documentation on Hugging Face
Ant Group chose to publish its training‑content summaries within the inclusionAI/AI‑Transparency repository hosted on Hugging Face. The repository’s description makes clear that it contains only documentation—no model weights or raw training datasets—and emphasizes that the act of publishing does not equate to any form of regulatory certification.
Even though the Ant Group documents conform to the Article 53(1)(d) template in terms of layout, they similarly leave the domain‑listing field empty. This mirrors the broader trend observed across the entire sample, reinforcing the notion that the omission is not isolated to a single provider.
Regulatory Context and Deadlines
The AI Act’s requirement to disclose scraped‑domain information becomes mandatory for newly marketed general‑purpose AI models as of August 2, 2025. For models that were already on the market before that date, a compliance window extends until August 2, 2027, giving existing providers a two‑year period to align their disclosures with the law.
Whether the current practice of using generic categories satisfies the literal legal wording of Article 53 remains a matter for the European Commission to decide. The Actuarial review carried out by ActuIA does not itself pronounce on the legality of the omissions.
- 21 summaries include the template but list no domains
- 3 DeepSeek summaries omit the section completely
- Providers use generic categories instead of specific URLs
- Ant Group’s Hugging Face repository clarifies that documents are not certification
For organisations operating in English‑speaking jurisdictions, the practical consequences are evident. Auditors conducting EU compliance checks are likely to flag the absence of explicit domain listings as an incomplete disclosure, which could trigger formal requests for additional information or corrective actions from European regulators.
Legal counsel advising AI developers should therefore prepare to augment existing summaries with concrete domain data. Should the Commission interpret the current generic descriptions as insufficient, providers may need to submit revised documents that enumerate the exact first‑ and second‑level domains that dominate their scraped data pools.
The broader industry implication is that the lack of specificity may hinder transparency objectives embedded in the AI Act. Stakeholders seeking to assess data provenance will find it more difficult to evaluate the risk profiles of models when the underlying sources are described only in vague terms.
Moreover, the pattern of omission could affect market perception. Companies that proactively disclose detailed domain lists may be viewed as more trustworthy by both regulators and end‑users, potentially gaining a competitive edge in a market increasingly focused on responsible AI practices.
Future monitoring efforts will need to track whether the European Commission issues formal guidance or enforcement notices concerning these omissions. Such guidance could clarify whether a simple categorical description satisfies the letter of the law or whether a precise enumeration is mandatory.
In summary, the review highlights a systemic shortfall in the way AI providers are handling the domain‑listing requirement of the AI Act. While the template is present in most documents, the substantive content is missing, raising questions about the overall effectiveness of current compliance strategies.
Localised note: In France, the Autorité de régulation des activités financières (ARAF) has indicated that it will closely monitor AI‑related disclosures, and French‑based AI firms are advised to prepare detailed domain inventories to avoid potential penalties under national enforcement mechanisms.
Sources
- Résumés AI Act : 21 documents sur 24 ne nomment aucun site moissonné dans la rubrique prévueActuIA · September 25, 2026
- Ant Ling — Training Content SummariesHugging Face · September 24, 2026



