# Model Documentation Evaluation Template # Metadata model_name: "Apertus" model_link: "https://huggingface.co/swiss-ai/Apertus-70B-2509" organization: "Swiss AI Initiative" org_link: "https://www.swiss-ai.org/" evaluation_date: "2026-03-30" public_summary_link: "https://huggingface.co/swiss-ai/Apertus-70B-2509/blob/main/Apertus_EU_Public_Summary.pdf" public_summary_date: "2025-09-01" public_summary_location: "https://huggingface.co/swiss-ai/Apertus-70B-2509" model_publication_date: "2025-09-02" category: "New model" # New model, fine-tuned model archive_file_name: "Apertus_EU_Public_Summary -- 2025_11_12.pdf" previous_versions: - "2026-01-12" # Requirement assessments # Each requirement should be assessed once with a score from 0-10 # Leave blank for N/A, and assign a max score of 0 S1: # Document-level requirements D1: # Clarity score: 9 max_score: 11 notes: "Document should clearly whether it is the latest version or if it is outdated and a replacement is made available. Document should provide a link to all versions of the document." D2: # Completeness score: 4 max_score: 4 notes: "" D3: # Consistency score: 7 max_score: 7 notes: "" D4: # Correctness score: 3.5 max_score: 5 notes: "Document should provide link to authoritative source of the document. Date format for date of last update should be correct." D5: # Accessibility score: 10 max_score: 10 notes: "" D6: # Comprehension score: 6 max_score: 9 notes: "Document should clearly indicate changes from previous version, as well as where notice of updates or changes will be provided" S2: # General information D1: # Clarity score: 20 max_score: 42 notes: "Description of types of content can be made more precise, i.e. categories of content such as legal text, social media comments rather than a generic description of 'public text-only data derived mainly from web documents'.

Our suggestion: specify categories in which the public datasets are actually provided, mentioning that training data primarily consists of educational webpages (from FineWeb-Edu and DCLM-Edu), crawled web-pages filtered for high-quality content (FineWeb-2, FineWeb-HQ, and FineWeb2-HQ), code data (The Stack dedup and StackV2_Edu_Filtered), and mathematical pretraining data (FineMath and MegaMath).

Mention of training data being fully transparent and reproducible is not strictly necessary given the field. We think this would be better in Additional comments (optional) field.

Other relevant characteristics of the overall training data: respecting of consent and the removal of toxic content, which should be provided in the assigned field in section 3.1 (consent), 3.2 (toxic content) or 3.3 (optional info)." D2: # Completeness score: 51 max_score: 62 notes: "Description of the linguistic characteristics of the overall training data: Missing info on language and EU languages.

Our suggestion: given the large number of languages used in pretraining (999+), mention the EU languages covered, e.g. 'text from all EU languages was used in the training of this model', with perhaps a link to the list of language-script pairs or where to find info on languages. This would already provide enough information to writers of a given language to be able to see whether training may have spanned their writing in that language." D3: # Consistency score: 8 max_score: 8 notes: "" D4: # Correctness score: 18 max_score: 18 notes: "" D5: # Accessibility score: 5 max_score: 5 notes: "" D6: # Comprehension score: 10 max_score: 11 notes: "Latest date of data acquisition/collection for model training: indicate whether continuous training is employed (see template placeholder text).

Our suggestion: Even though this is fairly evident given the type of model which Apertus is, it is still handy to include, and is expected by the template. The phrasing 'later date' is also ambiguous with fine-tuning." S3: # Public datasets D1: # Clarity score: 92 max_score: 92 notes: "" D2: # Completeness score: 95 max_score: 95 notes: "" D3: # Consistency score: 23 max_score: 23 notes: "" D4: # Correctness score: 23 max_score: 23 notes: "" D5: # Accessibility score: 20 max_score: 20 notes: "" D6: # Comprehension score: 19 max_score: 19 notes: "" S4: # Private datasets D1: # Clarity score: 0 max_score: 0 notes: "" D2: # Completeness score: 2 max_score: 2 notes: "" D3: # Consistency score: 3 max_score: 3 notes: "" D4: # Correctness score: 2 max_score: 2 notes: "" D5: # Accessibility score: 0 max_score: 0 notes: "" D6: # Comprehension score: 0 max_score: 0 notes: "" S5: # Scraped/crawled data D1: # Clarity score: 0 max_score: 0 notes: "" D2: # Completeness score: 3 max_score: 3 notes: "" D3: # Consistency score: 0 max_score: 0 notes: "" D4: # Correctness score: 3 max_score: 3 notes: "" D5: # Accessibility score: 0 max_score: 0 notes: "" D6: # Comprehension score: 0 max_score: 0 notes: "" S6: # User data D1: # Clarity score: 0 max_score: 0 notes: "" D2: # Completeness score: 8 max_score: 8 notes: "" D3: # Consistency score: 0 max_score: 0 notes: "" D4: # Correctness score: 8 max_score: 8 notes: "" D5: # Accessibility score: 0 max_score: 0 notes: "" D6: # Comprehension score: 0 max_score: 0 notes: "" S7: # Synthetic & other D1: # Clarity score: 0 max_score: 0 notes: "" D2: # Completeness score: 3 max_score: 3 notes: "" D3: # Consistency score: 6 max_score: 6 notes: "" D4: # Correctness score: 3 max_score: 3 notes: "" D5: # Accessibility score: 0 max_score: 0 notes: "" D6: # Comprehension score: 0 max_score: 0 notes: "" S8: # Data processing D1: # Clarity score: 24 max_score: 24 notes: "" D2: # Completeness score: 19 max_score: 19 notes: "" D3: # Consistency score: 3 max_score: 3 notes: "" D4: # Correctness score: 37 max_score: 37 notes: "" D5: # Accessibility score: 15 max_score: 15 notes: "" D6: # Comprehension score: 51 max_score: 51 notes: "" general_notes: "The public summary for the Apertus model family can be found as a PDF on HuggingFace in the same context as its models, for instance at the repository of Apertus-70B-2509. Each field of the template was filled in, including explicitly marking sections as not applicable, though some fields included superfluous information not relevant to the topic or question which caused a few points deduction. We assessed its score to be 92.90% with Grade A for transparency, and 97.14% with Grade A+ for usefulness, which were the highest of all assessed summaries published before and during our initial research."