# Model Documentation Evaluation Template # Metadata model_name: "Phi-4" model_link: "https://huggingface.co/microsoft/phi-4" organization: "Microsoft" org_link: "https://www.microsoft.com/en-us/research/" evaluation_date: "2026-01-12" public_summary_link: "https://huggingface.co/microsoft/phi-4/blob/main/data_summary_card.md" public_summary_date: "2025-11-24" public_summary_location: "https://huggingface.co/microsoft/phi-4" model_publication_date: "2024-12-12" category: "New model" # New model, fine-tuned model archive_file_name: "Phi_4_2026_02_05.pdf" # Requirement assessments # Each requirement should be assessed once with a score from 0-10 # Leave blank for N/A, and assign a max score of 0 S1: # Document-level requirements D1: # Clarity score: 9 max_score: 11 notes: "Document should clearly indicate whether it is the latest version or if it is outdated and a replacement is made available.

Document should provide link to all versions of the document." D2: # Completeness score: 4 max_score: 4 notes: "" D3: # Consistency score: 7 max_score: 7 notes: "" D4: # Correctness score: 3.5 max_score: 5 notes: "Document should provide link to authoritative source of the document. Date format should be correct." D5: # Accessibility score: 9.5 max_score: 10 notes: "Document must be easy to find." D6: # Comprehension score: 6 max_score: 9 notes: "Document should clearly indicate changes from previous version. Document should indicate where notice of updates or changes will be provided." S2: # General information D1: # Clarity score: 12 max_score: 24 notes: "Provider contact details refer to a generic contact point. Date of latest data acquisition should have the right format. Section 'Description of the linguistic characteristics of the overall training data' missing." D2: # Completeness score: 45 max_score: 57 notes: "Section 'Description of the linguistic characteristics of the overall training data' missing." D3: # Consistency score: 8 max_score: 8 notes: "" D4: # Correctness score: 7 max_score: 13 notes: "Section 'Description of the linguistic characteristics of the overall training data' missing." D5: # Accessibility score: 0 max_score: 0 notes: "" D6: # Comprehension score: 11 max_score: 11 notes: "" S3: # Public datasets D1: # Clarity score: 0 max_score: 52 notes: "No list of large publicly available datasets provided. No general description of other publicly available datasets provided." D2: # Completeness score: 1 max_score: 60 notes: "No modalities provided. No list of large publicly available datasets provided. No general description of other publicly available datasets provided." D3: # Consistency score: 3 max_score: 3 notes: "" D4: # Correctness score: 1 max_score: 20 notes: "No modalities provided. No general description of other publicly available datasets provided." D5: # Accessibility score: 0 max_score: 20 notes: "No list of large publicly available datasets provided." D6: # Comprehension score: 0 max_score: 16 notes: "No general description of other publicly available datasets provided." S4: # Private datasets D1: # Clarity score: 0 max_score: 0 notes: "" D2: # Completeness score: 2 max_score: 5 notes: "" D3: # Consistency score: 3 max_score: 3 notes: "" D4: # Correctness score: 2 max_score: 5 notes: "" D5: # Accessibility score: 0 max_score: 0 notes: "" D6: # Comprehension score: 0 max_score: 0 notes: "" S5: # Scraped/crawled data D1: # Clarity score: 0 max_score: 0 notes: "" D2: # Completeness score: 0 max_score: 3 notes: "Section missing." D3: # Consistency score: 0 max_score: 0 notes: "" D4: # Correctness score: 0 max_score: 3 notes: "Section missing." D5: # Accessibility score: 0 max_score: 0 notes: "" D6: # Comprehension score: 0 max_score: 0 notes: "" S6: # User data D1: # Clarity score: 0 max_score: 0 notes: "" D2: # Completeness score: 0 max_score: 8 notes: "Section missing except for generic statement." D3: # Consistency score: 0 max_score: 0 notes: "" D4: # Correctness score: 0 max_score: 8 notes: "Section missing except for generic statement." D5: # Accessibility score: 0 max_score: 0 notes: "" D6: # Comprehension score: 0 max_score: 0 notes: "" S7: # Synthetic & other D1: # Clarity score: 0 max_score: 24 notes: "Section only affirms that synthetic data was used." D2: # Completeness score: 1 max_score: 30 notes: "Section only affirms that synthetic data was used." D3: # Consistency score: 6 max_score: 6 notes: "" D4: # Correctness score: 1 max_score: 30 notes: "Section only affirms that synthetic data was used." D5: # Accessibility score: 0 max_score: 0 notes: "" D6: # Comprehension score: 0 max_score: 12 notes: "Section only affirms that synthetic data was used." S8: # Data processing D1: # Clarity score: 18 max_score: 18 notes: "" D2: # Completeness score: 0 max_score: 18 notes: "Section contains only generic statements and an affirmation that the data was cleaned." D3: # Consistency score: 3 max_score: 3 notes: "" D4: # Correctness score: 18 max_score: 36 notes: "Section contains only generic statements and an affirmation that the data was cleaned." D5: # Accessibility score: 0 max_score: 15 notes: "Section contains only generic statements and an affirmation that the data was cleaned." D6: # Comprehension score: 0 max_score: 15 notes: "Section contains only generic statements and an affirmation that the data was cleaned." general_notes: "We did not find an explicitly published summary for Microsoft's Phi-4 model, but found a file in its repository on HuggingFace called \"data_summary_card.md\", whose structure closely matched that of the template. The title used in the file was \"Data Summary\", which was also title used by HuggingFace in its SmolLM public summary. We are uncertain as to whether this represents an industry practice of co-opting of the formal title of \"public summary of training content\" present in the template with a broader and vague title that could be confused with other data-related documentation also provided with GPAI models. We opted to assess this model regardless based on the assumption of stakeholders viewing the document as fulfilling the template requirements, as explained earlier.

The assessment showed that fields provided in the document are filled in to some extent, but that the document itself is quite sparse and does not provide many details and also had sections missing. We further found that section numbers and questions had a significant mismatch from what was provided in the template. The document also suffered from subjective non-relevant statements, such as question 2.3.1 on whether public data was used to train the model being answered with \"Microsoft follows all relevant laws and regulations pertaining to personal information\" — which is neither relevant nor clear. We also found the document did not provide required information significantly, for example subsequent parts that enquire about the source and uses of personal data (2.4 User data in template) were found completely missing. Our assessment reflects these systematic major issues in the outcomes where the document scored 33.30% with a Grade D for transparency, and 24.54% with a Grade F for usefulness, which is the lowest score amongst all assessed summaries published before and during our initial research."