# Model Documentation Evaluation Template
# Metadata
model_name: "Apertus"
model_link: "https://huggingface.co/swiss-ai/Apertus-70B-2509"
organization: "Swiss AI Initiative"
org_link: "https://www.swiss-ai.org/"
evaluation_date: "2026-03-30"
public_summary_link: "https://huggingface.co/swiss-ai/Apertus-70B-2509/blob/main/Apertus_EU_Public_Summary.pdf"
public_summary_date: "2025-09-01"
public_summary_location: "https://huggingface.co/swiss-ai/Apertus-70B-2509"
model_publication_date: "2025-09-02"
category: "New model" # New model, fine-tuned model
archive_file_name: "Apertus_EU_Public_Summary -- 2025_11_12.pdf"
previous_versions:
- "2026-01-12"
# Requirement assessments
# Each requirement should be assessed once with a score from 0-10
# Leave blank for N/A, and assign a max score of 0
S1: # Document-level requirements
D1: # Clarity
score: 9
max_score: 11
notes: "Document should clearly whether it is the latest version or if it is outdated and a replacement is made available. Document should provide a link to all versions of the document."
D2: # Completeness
score: 4
max_score: 4
notes: ""
D3: # Consistency
score: 7
max_score: 7
notes: ""
D4: # Correctness
score: 3.5
max_score: 5
notes: "Document should provide link to authoritative source of the document. Date format for date of last update should be correct."
D5: # Accessibility
score: 10
max_score: 10
notes: ""
D6: # Comprehension
score: 6
max_score: 9
notes: "Document should clearly indicate changes from previous version, as well as where notice of updates or changes will be provided"
S2: # General information
D1: # Clarity
score: 20
max_score: 42
notes: "Description of types of content can be made more precise, i.e. categories of content such as legal text, social media comments rather than a generic description of 'public text-only data derived mainly from web documents'. \
\
Our suggestion: specify categories in which the public datasets are actually provided, mentioning that training data primarily consists of educational webpages (from FineWeb-Edu and DCLM-Edu), crawled web-pages filtered for high-quality content (FineWeb-2, FineWeb-HQ, and FineWeb2-HQ), code data (The Stack dedup and StackV2_Edu_Filtered), and mathematical pretraining data (FineMath and MegaMath). \
\
Mention of training data being fully transparent and reproducible is not strictly necessary given the field. We think this would be better in Additional comments (optional) field. \
\
Other relevant characteristics of the overall training data: respecting of consent and the removal of toxic content, which should be provided in the assigned field in section 3.1 (consent), 3.2 (toxic content) or 3.3 (optional info)."
D2: # Completeness
score: 51
max_score: 62
notes: "Description of the linguistic characteristics of the overall training data: Missing info on language and EU languages. \
\
Our suggestion: given the large number of languages used in pretraining (999+), mention the EU languages covered, e.g. 'text from all EU languages was used in the training of this model', with perhaps a link to the list of language-script pairs or where to find info on languages. This would already provide enough information to writers of a given language to be able to see whether training may have spanned their writing in that language."
D3: # Consistency
score: 8
max_score: 8
notes: ""
D4: # Correctness
score: 18
max_score: 18
notes: ""
D5: # Accessibility
score: 5
max_score: 5
notes: ""
D6: # Comprehension
score: 10
max_score: 11
notes: "Latest date of data acquisition/collection for model training: indicate whether continuous training is employed (see template placeholder text). \
\
Our suggestion: Even though this is fairly evident given the type of model which Apertus is, it is still handy to include, and is expected by the template. The phrasing 'later date' is also ambiguous with fine-tuning."
S3: # Public datasets
D1: # Clarity
score: 92
max_score: 92
notes: ""
D2: # Completeness
score: 95
max_score: 95
notes: ""
D3: # Consistency
score: 23
max_score: 23
notes: ""
D4: # Correctness
score: 23
max_score: 23
notes: ""
D5: # Accessibility
score: 20
max_score: 20
notes: ""
D6: # Comprehension
score: 19
max_score: 19
notes: ""
S4: # Private datasets
D1: # Clarity
score: 0
max_score: 0
notes: ""
D2: # Completeness
score: 2
max_score: 2
notes: ""
D3: # Consistency
score: 3
max_score: 3
notes: ""
D4: # Correctness
score: 2
max_score: 2
notes: ""
D5: # Accessibility
score: 0
max_score: 0
notes: ""
D6: # Comprehension
score: 0
max_score: 0
notes: ""
S5: # Scraped/crawled data
D1: # Clarity
score: 0
max_score: 0
notes: ""
D2: # Completeness
score: 3
max_score: 3
notes: ""
D3: # Consistency
score: 0
max_score: 0
notes: ""
D4: # Correctness
score: 3
max_score: 3
notes: ""
D5: # Accessibility
score: 0
max_score: 0
notes: ""
D6: # Comprehension
score: 0
max_score: 0
notes: ""
S6: # User data
D1: # Clarity
score: 0
max_score: 0
notes: ""
D2: # Completeness
score: 8
max_score: 8
notes: ""
D3: # Consistency
score: 0
max_score: 0
notes: ""
D4: # Correctness
score: 8
max_score: 8
notes: ""
D5: # Accessibility
score: 0
max_score: 0
notes: ""
D6: # Comprehension
score: 0
max_score: 0
notes: ""
S7: # Synthetic & other
D1: # Clarity
score: 0
max_score: 0
notes: ""
D2: # Completeness
score: 3
max_score: 3
notes: ""
D3: # Consistency
score: 6
max_score: 6
notes: ""
D4: # Correctness
score: 3
max_score: 3
notes: ""
D5: # Accessibility
score: 0
max_score: 0
notes: ""
D6: # Comprehension
score: 0
max_score: 0
notes: ""
S8: # Data processing
D1: # Clarity
score: 24
max_score: 24
notes: ""
D2: # Completeness
score: 19
max_score: 19
notes: ""
D3: # Consistency
score: 3
max_score: 3
notes: ""
D4: # Correctness
score: 37
max_score: 37
notes: ""
D5: # Accessibility
score: 15
max_score: 15
notes: ""
D6: # Comprehension
score: 51
max_score: 51
notes: ""
general_notes: "The public summary for the Apertus model family can be found as a PDF on HuggingFace in the same context as its models, for instance at the repository of Apertus-70B-2509. Each field of the template was filled in, including explicitly marking sections as not applicable, though some fields included superfluous information not relevant to the topic or question which caused a few points deduction. We assessed its score to be 92.90% with Grade A for transparency, and 97.14% with Grade A+ for usefulness, which were the highest of all assessed summaries published before and during our initial research."