# Model Documentation Evaluation Template
# Metadata
model_name: "Phi-4"
model_link: "https://huggingface.co/microsoft/phi-4"
organization: "Microsoft"
org_link: "https://www.microsoft.com/en-us/research/"
evaluation_date: "2026-01-12"
public_summary_link: "https://huggingface.co/microsoft/phi-4/blob/main/data_summary_card.md"
public_summary_date: "2025-11-24"
public_summary_location: "https://huggingface.co/microsoft/phi-4"
model_publication_date: "2024-12-12"
category: "New model" # New model, fine-tuned model
archive_file_name: "Phi_4_2026_02_05.pdf"
# Requirement assessments
# Each requirement should be assessed once with a score from 0-10
# Leave blank for N/A, and assign a max score of 0
S1: # Document-level requirements
D1: # Clarity
score: 9
max_score: 11
notes: "Document should clearly indicate whether it is the latest version or if it is outdated and a replacement is made available.
Document should provide link to all versions of the document."
D2: # Completeness
score: 4
max_score: 4
notes: ""
D3: # Consistency
score: 7
max_score: 7
notes: ""
D4: # Correctness
score: 3.5
max_score: 5
notes: "Document should provide link to authoritative source of the document. Date format should be correct."
D5: # Accessibility
score: 9.5
max_score: 10
notes: "Document must be easy to find."
D6: # Comprehension
score: 6
max_score: 9
notes: "Document should clearly indicate changes from previous version. Document should indicate where notice of updates or changes will be provided."
S2: # General information
D1: # Clarity
score: 12
max_score: 24
notes: "Provider contact details refer to a generic contact point. Date of latest data acquisition should have the right format. Section 'Description of the linguistic characteristics of the overall training data' missing."
D2: # Completeness
score: 45
max_score: 57
notes: "Section 'Description of the linguistic characteristics of the overall training data' missing."
D3: # Consistency
score: 8
max_score: 8
notes: ""
D4: # Correctness
score: 7
max_score: 13
notes: "Section 'Description of the linguistic characteristics of the overall training data' missing."
D5: # Accessibility
score: 0
max_score: 0
notes: ""
D6: # Comprehension
score: 11
max_score: 11
notes: ""
S3: # Public datasets
D1: # Clarity
score: 0
max_score: 52
notes: "No list of large publicly available datasets provided. No general description of other publicly available datasets provided."
D2: # Completeness
score: 1
max_score: 60
notes: "No modalities provided. No list of large publicly available datasets provided. No general description of other publicly available datasets provided."
D3: # Consistency
score: 3
max_score: 3
notes: ""
D4: # Correctness
score: 1
max_score: 20
notes: "No modalities provided. No general description of other publicly available datasets provided."
D5: # Accessibility
score: 0
max_score: 20
notes: "No list of large publicly available datasets provided."
D6: # Comprehension
score: 0
max_score: 16
notes: "No general description of other publicly available datasets provided."
S4: # Private datasets
D1: # Clarity
score: 0
max_score: 0
notes: ""
D2: # Completeness
score: 2
max_score: 5
notes: ""
D3: # Consistency
score: 3
max_score: 3
notes: ""
D4: # Correctness
score: 2
max_score: 5
notes: ""
D5: # Accessibility
score: 0
max_score: 0
notes: ""
D6: # Comprehension
score: 0
max_score: 0
notes: ""
S5: # Scraped/crawled data
D1: # Clarity
score: 0
max_score: 0
notes: ""
D2: # Completeness
score: 0
max_score: 3
notes: "Section missing."
D3: # Consistency
score: 0
max_score: 0
notes: ""
D4: # Correctness
score: 0
max_score: 3
notes: "Section missing."
D5: # Accessibility
score: 0
max_score: 0
notes: ""
D6: # Comprehension
score: 0
max_score: 0
notes: ""
S6: # User data
D1: # Clarity
score: 0
max_score: 0
notes: ""
D2: # Completeness
score: 0
max_score: 8
notes: "Section missing except for generic statement."
D3: # Consistency
score: 0
max_score: 0
notes: ""
D4: # Correctness
score: 0
max_score: 8
notes: "Section missing except for generic statement."
D5: # Accessibility
score: 0
max_score: 0
notes: ""
D6: # Comprehension
score: 0
max_score: 0
notes: ""
S7: # Synthetic & other
D1: # Clarity
score: 0
max_score: 24
notes: "Section only affirms that synthetic data was used."
D2: # Completeness
score: 1
max_score: 30
notes: "Section only affirms that synthetic data was used."
D3: # Consistency
score: 6
max_score: 6
notes: ""
D4: # Correctness
score: 1
max_score: 30
notes: "Section only affirms that synthetic data was used."
D5: # Accessibility
score: 0
max_score: 0
notes: ""
D6: # Comprehension
score: 0
max_score: 12
notes: "Section only affirms that synthetic data was used."
S8: # Data processing
D1: # Clarity
score: 18
max_score: 18
notes: ""
D2: # Completeness
score: 0
max_score: 18
notes: "Section contains only generic statements and an affirmation that the data was cleaned."
D3: # Consistency
score: 3
max_score: 3
notes: ""
D4: # Correctness
score: 18
max_score: 36
notes: "Section contains only generic statements and an affirmation that the data was cleaned."
D5: # Accessibility
score: 0
max_score: 15
notes: "Section contains only generic statements and an affirmation that the data was cleaned."
D6: # Comprehension
score: 0
max_score: 15
notes: "Section contains only generic statements and an affirmation that the data was cleaned."
general_notes: "We did not find an explicitly published summary for Microsoft's Phi-4 model, but found a file in its repository on HuggingFace called \"data_summary_card.md\", whose structure closely matched that of the template. The title used in the file was \"Data Summary\", which was also title used by HuggingFace in its SmolLM public summary. We are uncertain as to whether this represents an industry practice of co-opting of the formal title of \"public summary of training content\" present in the template with a broader and vague title that could be confused with other data-related documentation also provided with GPAI models. We opted to assess this model regardless based on the assumption of stakeholders viewing the document as fulfilling the template requirements, as explained earlier.
The assessment showed that fields provided in the document are filled in to some extent, but that the document itself is quite sparse and does not provide many details and also had sections missing. We further found that section numbers and questions had a significant mismatch from what was provided in the template. The document also suffered from subjective non-relevant statements, such as question 2.3.1 on whether public data was used to train the model being answered with \"Microsoft follows all relevant laws and regulations pertaining to personal information\" — which is neither relevant nor clear. We also found the document did not provide required information significantly, for example subsequent parts that enquire about the source and uses of personal data (2.4 User data in template) were found completely missing. Our assessment reflects these systematic major issues in the outcomes where the document scored 33.30% with a Grade D for transparency, and 24.54% with a Grade F for usefulness, which is the lowest score amongst all assessed summaries published before and during our initial research."