Overview of Quality Assessment Findings

Below is an overview of our quality assessment for each public summary. We scored each section of the summary based on our developed methodology, and then calculated the overall score and grade. Each summary is scored for two dimensions: Transparency (T) on how the information is provided, and Usefulness (U) for whether it is sufficient for stakeholders' needs. For convenience, all scores are reflected as a percentage (out of hundred) and the grades are expressed on a scale from A+ (highest) to F (lowest). Public summaries that were missing are marked as "N/A".

Section→ Grade Overall General Information Public Data Sources Private Data Sources Scraped/Crawled Data User Data Synthetic & Other Data Data Processing Document
Model↓ T U T U T U T U T U T U T U T U T U T U
Apertus A   A+ 92 97 74 93 100 100 100 N/A 100 N/A 100 N/A 100 N/A 100 100 87 84
Bria 3.2 B+ A   86 94 69 92 100 N/A 100 100 100 N/A 100 N/A 100 N/A 100 100 70 68
SmolLM3-3B B+ B+ 82 86 73 90 96 100 100 N/A 100 N/A 100 N/A 60 0 92 100 68 58
Bielik v3 11B Instruct B+ C+ 88 71 87 93 86 100 100 N/A 83 51 100 N/A 100 100 98 71 88 84
Phi-4 D   F   33 24 70 100 3 0 53 N/A 0 N/A 0 N/A 8 0 52 0 87 81
Claude Sonnet 4.5 N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A
Gemini 2.5 Flash Image N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A
GPT-5 N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A
GPT-OSS N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A
Sora 2 N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A N/A


Analysis based on following public summaries: