# Benchmark Dataset Inventory This file is generated from the current CLI registry and loader metadata. It is the public source-of-truth for what the repository supports at release time. - CLI benchmark registrations: **166** - Canonical benchmark entries listed below: **155** - Deprecated compatibility aliases: **11** - CLI modes: **4** Count semantics: counts are the default loader scope where the code pins one; otherwise counts use registry metadata when present. `-` means the loader follows the official upstream split but the repository does not pin a static count, so users should inspect the current dataset card or run a source audit in their environment. ## Registered Benchmarks All CLI-exposed canonical benchmark entries are listed in one table below. `hf_*` rows use the generic HuggingFace loader shown in the Loader column; non-`hf_*` rows use dedicated benchmark loaders. Training-only corpora and removed non-benchmark rows are intentionally excluded from this public inventory. | Benchmark | Source | Config / split | Count | Domain / content | Input | Task type | Answer/scorer | Gated | Network | Offline cache | Multimodal | Loader | | --- | --- | --- | ---: | --- | --- | --- | --- | --- | --- | --- | --- | --- | | `aa_lcr` | ArtificialAnalysis/AA-LCR | custom loader | upstream split size | long-context reasoning over documents | text, optional retrieved documents | long-context QA | openText | no | yes on first run | yes | no | `load_aa_lcr_tasks` | | `agentclinic` | AgentClinic official release | custom loader | loader default | doctor-patient diagnostic scenarios | text dialogue | clinical simulation | openText | no | source-dependent | yes | no | `load_agentclinic_tasks` | | `bioasq` | BioASQ official/local source | custom loader | loader default | factoid/list/yes-no biomedical questions | text | biomedical QA | openText | no | source-dependent | yes | no | `load_bioasq_tasks` | | `bioprobench` | BioProBench official data | custom loader | official benchmark scope | biological protocol understanding and repair | text | protocol QA | mixed | no | yes on first run | yes | no | `load_bioprobench_tasks` | | `bixbench` | futurehouse/BixBench | custom loader | 205 | biomedical data-analysis tasks with optional official data capsules | text + optional CapsuleFolder zip | bioinformatics agent QA | MCQ adaptation or open-answer official-compatible scorer | no | yes | yes | no | `load_bixbench_tasks`; `prepare-bixbench` for capsules | | `genotex` | GenoTEX official data | custom loader | loader default | genomics text reasoning | text/genomics | genomics QA | openText | no | source-dependent | yes | no | `load_genotex_tasks` | | `gpqa_bio` | Idavidrein/gpqa:gpqa_diamond/train | custom loader | 198 | graduate-level biology/chemistry/medicine questions | text | graduate science MCQ | multipleChoice | yes | yes | yes | no | `load_gpqa_bio_tasks` | | `healthbench` | OpenAI HealthBench official data | custom loader | official benchmark scope | consumer-health answer quality and safety | text conversation | health conversation | openText | source-dependent | yes on first run | yes | no | `load_healthbench_tasks` | | `hle_gold` | futurehouse/hle-gold-bio-chem:train | custom loader | 149 | HLE Gold bio/chem subset | text | expert QA | mixed | yes | yes | yes | no | `load_hle_gold_tasks` | | `labbench` | futurehouse/lab-bench | custom loader | loader default subsets | LitQA, cloning, protocol tasks | text | biomedical agent QA | mixed | yes | yes | yes | no | `load_labbench_tasks` | | `labbench2` | EdisonScientific/labbench2 text-only subsets | custom loader | 821 | LAB-Bench 2 text-only evaluation subset | text | literature/database/patent QA | openText | yes | yes | yes | no by default | `load_labbench2_tasks` | | `medagentbench` | MedAgentBench official data | custom loader | loader default | clinical workflow and EHR tasks | text/EHR | medical agent workflow | mixed | no | source-dependent | yes | no | `load_medagentbench_tasks` | | `medcalc` | ncbi/MedCalc-Bench-v1.2:test | custom loader | 1100 | medical calculator word problems | text | clinical calculation | exactNumeric | no | yes | yes | no | `load_medcalc_tasks` | | `medhelm` | MedHELM official/public sources | custom loader | official benchmark scope | medical QA, safety, and scenario tasks | text | medical HELM tasks | mixed | source-dependent | yes on first run | yes | no | `load_medhelm_tasks` | | `medmcqa` | openlifescienceai/medmcqa | custom loader | loader default split | medical entrance-exam questions | text | medical MCQ | multipleChoice | no | yes | yes | no | `load_medical_qa_tasks` | | `medqa` | GBaker/MedQA-USMLE-4-options | custom loader | loader default split | USMLE-style questions | text | USMLE MCQ | multipleChoice | no | yes | yes | no | `load_medical_qa_tasks` | | `medxpertqa` | TsinghuaC3I/MedXpertQA | custom loader | loader default Text subset | expert medical reasoning questions | text | expert medical MCQ | multipleChoice | no | yes | yes | no | `load_medxpertqa_tasks` | | `medxpertqa_mm` | TsinghuaC3I/MedXpertQA-MM | custom loader | official multimodal subset | expert medical multimodal questions | text with optional images | medical VQA/MCQ | multipleChoice | no | yes | yes | yes; text fallback by default | `load_medxpertqa_mm_tasks` | | `mmlu` | MMLU medical/biology subjects | custom loader | loader default subjects | MMLU anatomy, medicine, biology, genetics subjects | text | academic MCQ | multipleChoice | no | yes | yes | no | `load_mmlu_tasks` | | `pathvqa` | PathVQA official/HF source | custom loader | loader default split | pathology image questions | text+image | pathology VQA | openText | no | yes | yes | yes | `load_pathvqa_tasks` | | `pubmedqa` | qiaojin/PubMedQA or OpenLifeScience mirror | custom loader | loader default split | yes/no/maybe biomedical literature questions | text abstract | PubMed abstract QA | multipleChoice | no | yes | yes | no | `load_medical_qa_tasks` | | `quick_suite` | built-in repository fixtures | custom loader | 20 | 5 MCQ, 5 exact, 5 numeric, 5 open-text scorer checks | text | offline smoke | mixed | no | no | not needed | no | `load_quick_suite_tasks` | | `rag_essential` | built-in RAG essential tasks | custom loader | 12 | tasks designed to reward retrieval/tool use | text | retrieval/tool-use QA | openText | no | no | not needed | no | `load_rag_essential_tasks` | | `super_chemistry` | ZehuaZhao/SUPERChem:SUPERChem-500.parquet | custom loader | 500 text rows by default; 500 official rows total | advanced chemistry questions | text, optional images | chemistry MCQ | multipleChoice | no | yes | yes | yes; text fallback by default | `load_super_chemistry_tasks` | | `superchem` | SuperChem official data | custom loader | loader default | chemistry evaluation tasks | text | chemistry QA | mixed | source-dependent | yes on first run | yes | no | `load_superchem_tasks` | | `supergpqa` | SuperGPQA official data | custom loader | loader default | graduate-level science questions | text | science MCQ | multipleChoice | source-dependent | yes on first run | yes | no | `load_supergpqa_tasks` | | `hf_adaptllm_chemprot` | [`AdaptLLM/medicine-tasks`](https://huggingface.co/datasets/AdaptLLM/medicine-tasks) | config: ChemProt; split: default | 500 (test) | biomedical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_adaptllm_medicine_tasks` | [`AdaptLLM/medicine-tasks`](https://huggingface.co/datasets/AdaptLLM/medicine-tasks) | config: USMLE; split: default | 1273 (test) | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_adaptllm_mqp` | [`AdaptLLM/medicine-tasks`](https://huggingface.co/datasets/AdaptLLM/medicine-tasks) | config: MQP; split: default | 610 (test) | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_adaptllm_rct` | [`AdaptLLM/medicine-tasks`](https://huggingface.co/datasets/AdaptLLM/medicine-tasks) | config: RCT; split: default | 2000 (test) | biomedical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_ade_corpus_v2` | [`ade-benchmark-corpus/ade_corpus_v2`](https://huggingface.co/datasets/ade-benchmark-corpus/ade_corpus_v2) | config: Ade_corpus_v2_classification; split: train | 23516 | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_anatem` | [`bigbio/anat_em`](https://huggingface.co/datasets/bigbio/anat_em) | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_bacbench_antibiotic_resistance_dna` | [`macwiatrak/bacbench-antibiotic-resistance-dna`](https://huggingface.co/datasets/macwiatrak/bacbench-antibiotic-resistance-dna) | split: default | - | dna | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_bacbench_phenotypic_traits_dna` | [`macwiatrak/bacbench-phenotypic-traits-dna`](https://huggingface.co/datasets/macwiatrak/bacbench-phenotypic-traits-dna) | split: default | - | dna | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_bc2gm` | [`spyysalo/bc2gm_corpus`](https://huggingface.co/datasets/spyysalo/bc2gm_corpus) | config: bc2gm_corpus; split: test | 5000 | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_bc5cdr` | [`EMBO/BLURB`](https://huggingface.co/datasets/EMBO/BLURB) | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_bigbio_med_qa` | [`bigbio/med_qa`](https://huggingface.co/datasets/bigbio/med_qa) | split: default | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_bigbio_pubmed_qa` | [`bigbio/pubmed_qa`](https://huggingface.co/datasets/bigbio/pubmed_qa) | split: default | - | medical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_biocreative_viii_biored` | [`bigbio/biored`](https://huggingface.co/datasets/bigbio/biored) | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_biomedbench` | [`biomedbench/BioMedBench`](https://huggingface.co/datasets/biomedbench/BioMedBench) | split: default | - | biomedical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_biored` | [`bigbio/biored`](https://huggingface.co/datasets/bigbio/biored) | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_biosses` | [`mteb/biosses-sts`](https://huggingface.co/datasets/mteb/biosses-sts) | split: test | - | biomedical | text/structured | regression | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_blurb` | [`EMBO/BLURB`](https://huggingface.co/datasets/EMBO/BLURB) | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_careqa` | [`HPAI-BSC/CareQA`](https://huggingface.co/datasets/HPAI-BSC/CareQA) | config: CareQA_en; split: test | 5621 | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_ccdv_pubmed_summarization` | [`ccdv/pubmed-summarization`](https://huggingface.co/datasets/ccdv/pubmed-summarization) | split: default | - | biomedical | text/structured | summarization | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_chembench` | [`jablonkagroup/ChemBench`](https://huggingface.co/datasets/jablonkagroup/ChemBench) | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_chemistry_qa` | [`avaliev/ChemistryQA`](https://huggingface.co/datasets/avaliev/ChemistryQA) | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_chemllmbench` | [`blc-org/chemllmbench`](https://huggingface.co/datasets/blc-org/chemllmbench) | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_clicr` | [`bigbio/clicr`](https://huggingface.co/datasets/bigbio/clicr) | split: default | - | clinical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_clinical_trials_eligibility_nlp` | [`bigbio/n2c2_2018_track1`](https://huggingface.co/datasets/bigbio/n2c2_2018_track1) | split: default | - | clinical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_cmb` | [`FreedomIntelligence/CMB`](https://huggingface.co/datasets/FreedomIntelligence/CMB) | config: CMB-Exam; split: test | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_cmexam` | [`fzkuji/CMExam`](https://huggingface.co/datasets/fzkuji/CMExam) | split: test | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_cord19_qa` | [`allenai/cord19`](https://huggingface.co/datasets/allenai/cord19) | split: default | - | biomedical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_craft` | [`bigbio/craft`](https://huggingface.co/datasets/bigbio/craft) | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_ddi_corpus_2013` | [`OpenMed/DDI-Corpus-Processed`](https://huggingface.co/datasets/OpenMed/DDI-Corpus-Processed) | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_discoverybench_biomedical` | [`allenai/discoverybench`](https://huggingface.co/datasets/allenai/discoverybench) | split: train | - | biomedical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_ebm_nlp` | [`bigbio/ebm_pico`](https://huggingface.co/datasets/bigbio/ebm_pico) | config: processed; split: test | - | biomedical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_evidence_inference` | [`hpi-dhc/evidence-inference-simple`](https://huggingface.co/datasets/hpi-dhc/evidence-inference-simple) | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_fgbench` | [`xuan-liu/FGBench`](https://huggingface.co/datasets/xuan-liu/FGBench) | split: test | - | chemistry | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_fluorescence_prediction` | [`proteinglm/fluorescence_prediction`](https://huggingface.co/datasets/proteinglm/fluorescence_prediction) | split: default | - | protein | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_gad` | [`bigbio/gad`](https://huggingface.co/datasets/bigbio/gad) | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_gaianet_chemistry` | [`gaianet/chemistry`](https://huggingface.co/datasets/gaianet/chemistry) | split: default | - | chemistry | text/structured | text | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_genbio_proteingym_dms` | [`genbio-ai/ProteinGYM-DMS`](https://huggingface.co/datasets/genbio-ai/ProteinGYM-DMS) | split: default | - | protein | text/structured | protein_fitness | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_geneturing` | [`vladimire/geneturing`](https://huggingface.co/datasets/vladimire/geneturing) | config: all; split: test | 600 | genomics | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_genomics_long_range` | [`InstaDeepAI/genomics-long-range-benchmark`](https://huggingface.co/datasets/InstaDeepAI/genomics-long-range-benchmark) | split: default | - | dna | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_hallmarks_of_cancer` | [`bigbio/hallmarks_of_cancer`](https://huggingface.co/datasets/bigbio/hallmarks_of_cancer) | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_headqa` | [`openlifescienceai/headqa`](https://huggingface.co/datasets/openlifescienceai/headqa) | split: test | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_healthqa` | [`nlplabtdtu/health_qa`](https://huggingface.co/datasets/nlplabtdtu/health_qa) | split: default | - | medical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_icml2022_proteingym` | [`ICML2022/ProteinGym`](https://huggingface.co/datasets/ICML2022/ProteinGym) | split: default | - | protein | text/structured | protein_fitness | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_jnlpba` | [`EMBO/BLURB`](https://huggingface.co/datasets/EMBO/BLURB) | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_katielink_moleculenet_bace` | [`katielink/moleculenet-benchmark`](https://huggingface.co/datasets/katielink/moleculenet-benchmark) | config: bace; split: default | 152 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_katielink_moleculenet_bbbp` | [`katielink/moleculenet-benchmark`](https://huggingface.co/datasets/katielink/moleculenet-benchmark) | config: bbbp; split: default | 194 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_katielink_moleculenet_clintox` | [`katielink/moleculenet-benchmark`](https://huggingface.co/datasets/katielink/moleculenet-benchmark) | config: clintox; split: default | 143 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_katielink_moleculenet_esol` | [`katielink/moleculenet-benchmark`](https://huggingface.co/datasets/katielink/moleculenet-benchmark) | config: esol; split: default | 113 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_katielink_moleculenet_freesolv` | [`katielink/moleculenet-benchmark`](https://huggingface.co/datasets/katielink/moleculenet-benchmark) | config: freesolv; split: default | 65 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_katielink_moleculenet_hiv` | [`katielink/moleculenet-benchmark`](https://huggingface.co/datasets/katielink/moleculenet-benchmark) | config: hiv; split: default | 4113 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_katielink_moleculenet_lipo` | [`katielink/moleculenet-benchmark`](https://huggingface.co/datasets/katielink/moleculenet-benchmark) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_katielink_moleculenet_sider` | [`katielink/moleculenet-benchmark`](https://huggingface.co/datasets/katielink/moleculenet-benchmark) | config: sider; split: default | 143 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_katielink_moleculenet_tox21` | [`katielink/moleculenet-benchmark`](https://huggingface.co/datasets/katielink/moleculenet-benchmark) | config: tox21; split: default | 783 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_litcovid` | [`ncats/litcovid`](https://huggingface.co/datasets/ncats/litcovid) | split: validation | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_liveqa_med` | [`hyesunyun/liveqa_medical_trec2017`](https://huggingface.co/datasets/hyesunyun/liveqa_medical_trec2017) | split: test | - | medical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_longhealth` | [`tonychenxyz/longhealth`](https://huggingface.co/datasets/tonychenxyz/longhealth) | config: plain; split: test | 400 | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_lpm24_eval_caption` | [`language-plus-molecules/LPM-24_eval-caption`](https://huggingface.co/datasets/language-plus-molecules/LPM-24_eval-caption) | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_lpm24_eval_molgen` | [`language-plus-molecules/LPM-24_eval-molgen`](https://huggingface.co/datasets/language-plus-molecules/LPM-24_eval-molgen) | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_medcase_reasoning` | [`zou-lab/MedCaseReasoning`](https://huggingface.co/datasets/zou-lab/MedCaseReasoning) | split: test | - | clinical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_medconceptsqa` | [`ofir408/MedConceptsQA`](https://huggingface.co/datasets/ofir408/MedConceptsQA) | config: all; split: default | 819772 (test) | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_meddialogqa` | [`UCSD26/medical_dialog`](https://huggingface.co/datasets/UCSD26/medical_dialog) | split: default | - | medical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_medexqa` | [`bluesky333/MedExQA`](https://huggingface.co/datasets/bluesky333/MedExQA) | split: default | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_medical_question_pairs` | [`curaihealth/medical_questions_pairs`](https://huggingface.co/datasets/curaihealth/medical_questions_pairs) | split: default | - | medical | text/structured | pair_classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_medication_qa` | [`truehealth/medicationqa`](https://huggingface.co/datasets/truehealth/medicationqa) | split: train | - | medical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_medmcqa_explanations` | [`openlifescienceai/medmcqa`](https://huggingface.co/datasets/openlifescienceai/medmcqa) | split: validation | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_mednli` | [`araag2/MedNLI`](https://huggingface.co/datasets/araag2/MedNLI) | config: processed; split: test | 1422 | clinical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_medpalm_eval_set` | [`katielink/healthsearchqa`](https://huggingface.co/datasets/katielink/healthsearchqa) | config: 140_question_subset; split: train | 140 | medical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_medpub_qa` | [`qiaojin/PubMedQA`](https://huggingface.co/datasets/qiaojin/PubMedQA) | config: pqa_labeled; split: train | 1000 | biomedical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_medqa_taiwan` | [`xuxuxuxuxu/MedQA_Taiwan_test`](https://huggingface.co/datasets/xuxuxuxuxu/MedQA_Taiwan_test) | split: default | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_medquad` | [`keivalya/MedQuad-MedicalQnADataset`](https://huggingface.co/datasets/keivalya/MedQuad-MedicalQnADataset) | split: default | - | medical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_meds_bench` | [`Henrychur/MedS-Bench`](https://huggingface.co/datasets/Henrychur/MedS-Bench) | split: default | - | medical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_meqsum` | [`albertvillanova/meqsum`](https://huggingface.co/datasets/albertvillanova/meqsum) | split: train | - | medical | text/structured | summarization | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_mol_instructions_pubchemqa` | [`zjunlp/Mol-Instructions`](https://huggingface.co/datasets/zjunlp/Mol-Instructions) | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculeace` | [`karina-zadorozhny/moleculeace`](https://huggingface.co/datasets/karina-zadorozhny/moleculeace) | config: CHEMBL1862_Ki; split: default | 161 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculeace_chembl1871_ki` | [`karina-zadorozhny/moleculeace`](https://huggingface.co/datasets/karina-zadorozhny/moleculeace) | config: CHEMBL1871_Ki; split: default | 134 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculeace_chembl204_ki` | [`karina-zadorozhny/moleculeace`](https://huggingface.co/datasets/karina-zadorozhny/moleculeace) | config: CHEMBL204_Ki; split: default | 553 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculeace_chembl214_ki` | [`karina-zadorozhny/moleculeace`](https://huggingface.co/datasets/karina-zadorozhny/moleculeace) | config: CHEMBL214_Ki; split: default | 666 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculeace_chembl228_ki` | [`karina-zadorozhny/moleculeace`](https://huggingface.co/datasets/karina-zadorozhny/moleculeace) | config: CHEMBL228_Ki; split: default | 342 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculeace_chembl237_ec50` | [`karina-zadorozhny/moleculeace`](https://huggingface.co/datasets/karina-zadorozhny/moleculeace) | config: CHEMBL237_EC50; split: default | 193 (test) | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculenet_bace` | [`scikit-fingerprints/MoleculeNet_BACE`](https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_BACE) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculenet_bbbp` | [`scikit-fingerprints/MoleculeNet_BBBP`](https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_BBBP) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculenet_clintox` | [`scikit-fingerprints/MoleculeNet_ClinTox`](https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ClinTox) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculenet_esol` | [`scikit-fingerprints/MoleculeNet_ESOL`](https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ESOL) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculenet_freesolv` | [`scikit-fingerprints/MoleculeNet_FreeSolv`](https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_FreeSolv) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculenet_hiv` | [`scikit-fingerprints/MoleculeNet_HIV`](https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_HIV) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculenet_lipophilicity` | [`scikit-fingerprints/MoleculeNet_Lipophilicity`](https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_Lipophilicity) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculenet_pcba` | [`scikit-fingerprints/MoleculeNet_PCBA`](https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_PCBA) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculenet_sider` | [`scikit-fingerprints/MoleculeNet_SIDER`](https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_SIDER) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_moleculenet_toxcast` | [`scikit-fingerprints/MoleculeNet_ToxCast`](https://huggingface.co/datasets/scikit-fingerprints/MoleculeNet_ToxCast) | split: default | - | chemistry | text/structured | molecule_property | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_mollangbench` | [`ChemFM/MolLangBench`](https://huggingface.co/datasets/ChemFM/MolLangBench) | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_ms2` | [`allenai/mslr2022`](https://huggingface.co/datasets/allenai/mslr2022) | split: validation | - | biomedical | text/structured | summarization | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_mteb_medical_qa` | [`mteb/medical_qa`](https://huggingface.co/datasets/mteb/medical_qa) | split: default | - | medical | text/structured | retrieval | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_mteb_medical_retrieval` | [`mteb/MedicalRetrieval`](https://huggingface.co/datasets/mteb/MedicalRetrieval) | split: default | - | medical | text/structured | retrieval | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_mts_dialogue_clinical_note` | [`har1/MTS_Dialogue-Clinical_Note`](https://huggingface.co/datasets/har1/MTS_Dialogue-Clinical_Note) | split: default | - | clinical | text/structured | summarization | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_ncbi_disease` | [`EMBO/BLURB`](https://huggingface.co/datasets/EMBO/BLURB) | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_nlmchem` | [`jablonkagroup/nlmchem`](https://huggingface.co/datasets/jablonkagroup/nlmchem) | config: instruction_0; split: test | 404 | biomedical | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_pgr` | [`lasigeBioTM/PGR`](https://huggingface.co/datasets/lasigeBioTM/PGR) | split: test | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_ppi_benchmark` | [`bigbio/bioinfer`](https://huggingface.co/datasets/bigbio/bioinfer) | split: test | - | protein | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_protein_binding_sequences` | [`ronig/protein_binding_sequences`](https://huggingface.co/datasets/ronig/protein_binding_sequences) | split: default | - | protein | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_protein_deeploc` | [`proteinea/deeploc`](https://huggingface.co/datasets/proteinea/deeploc) | split: default | - | protein | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_protein_fluorescence` | [`proteinea/fluorescence`](https://huggingface.co/datasets/proteinea/fluorescence) | split: default | - | protein | text/structured | regression | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_protein_secondary_structure` | [`lamm-mit/protein_secondary_structure_from_PDB`](https://huggingface.co/datasets/lamm-mit/protein_secondary_structure_from_PDB) | split: default | - | protein | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_protein_solubility` | [`proteinea/solubility`](https://huggingface.co/datasets/proteinea/solubility) | split: default | - | protein | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_protein_stability` | [`SaProtHub/Dataset-Meta-scale-protein-stability`](https://huggingface.co/datasets/SaProtHub/Dataset-Meta-scale-protein-stability) | split: default | - | protein | text/structured | regression | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_proteingym_v01` | [`OATML-Markslab/ProteinGym_v0.1`](https://huggingface.co/datasets/OATML-Markslab/ProteinGym_v0.1) | split: default | - | protein | text/structured | protein_fitness | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_proteingym_v1` | [`OATML-Markslab/ProteinGym_v1`](https://huggingface.co/datasets/OATML-Markslab/ProteinGym_v1) | split: default | - | protein | text/structured | protein_fitness | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_proteinlmbench` | [`tsynbio/ProteinLMBench`](https://huggingface.co/datasets/tsynbio/ProteinLMBench) | config: evaluation; split: train | 944 | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_proteinlmbench_enzyme_cot` | [`tsynbio/ProteinLMBench`](https://huggingface.co/datasets/tsynbio/ProteinLMBench) | config: Enzyme_CoT; split: default | 10826 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_proteinlmbench_uniprot_disease` | [`tsynbio/ProteinLMBench`](https://huggingface.co/datasets/tsynbio/ProteinLMBench) | config: UniProt_Involvement in disease; split: default | 5575 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_proteinlmbench_uniprot_function` | [`tsynbio/ProteinLMBench`](https://huggingface.co/datasets/tsynbio/ProteinLMBench) | config: UniProt_Function; split: default | 464737 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_proteinlmbench_uniprot_induction` | [`tsynbio/ProteinLMBench`](https://huggingface.co/datasets/tsynbio/ProteinLMBench) | config: UniProt_Induction; split: default | 25359 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_proteinlmbench_uniprot_ptm` | [`tsynbio/ProteinLMBench`](https://huggingface.co/datasets/tsynbio/ProteinLMBench) | config: UniProt_Post-translational modification; split: default | 45783 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_proteinlmbench_uniprot_subunit` | [`tsynbio/ProteinLMBench`](https://huggingface.co/datasets/tsynbio/ProteinLMBench) | config: UniProt_Subunit structure; split: default | 291467 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_proteinlmbench_uniprot_tissue` | [`tsynbio/ProteinLMBench`](https://huggingface.co/datasets/tsynbio/ProteinLMBench) | config: UniProt_Tissue specificity; split: default | 50316 (train) | protein | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_pubmed_200k_rct` | [`pietrolesci/pubmed-200k-rct`](https://huggingface.co/datasets/pietrolesci/pubmed-200k-rct) | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_pubmed_abstract_classification` | [`uiyunkim-hub/pubmed-abstract`](https://huggingface.co/datasets/uiyunkim-hub/pubmed-abstract) | split: default | - | biomedical | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_raredis` | [`guan-wang/ReDis-QA`](https://huggingface.co/datasets/guan-wang/ReDis-QA) | split: test | - | medical | text/structured | mcq | multipleChoice | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_rna_downstream_tasks` | [`genbio-ai/rna-downstream-tasks`](https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks) | config: modification_site; split: test | 1200 | rna | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_rna_expression_hek` | [`genbio-ai/rna-downstream-tasks`](https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks) | config: expression_HEK; split: default | 14410 (train) | rna | text/structured | regression | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_rna_expression_muscle` | [`genbio-ai/rna-downstream-tasks`](https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks) | config: expression_Muscle; split: default | 1257 (train) | rna | text/structured | regression | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_rna_expression_pc3` | [`genbio-ai/rna-downstream-tasks`](https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks) | config: expression_pc3; split: default | 12579 (train) | rna | text/structured | regression | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_rna_mean_ribosome_load` | [`genbio-ai/rna-downstream-tasks`](https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks) | config: mean_ribosome_load; split: default | 7600 (test) | rna | text/structured | regression | exactNumeric | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_rna_modification_site` | [`genbio-ai/rna-downstream-tasks`](https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks) | config: modification_site; split: default | 1200 (test) | rna | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_rna_ncrna_family_bnoise0` | [`genbio-ai/rna-downstream-tasks`](https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks) | config: ncrna_family_bnoise0; split: default | 25342 (test) | rna | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_rna_splice_site_acceptor` | [`genbio-ai/rna-downstream-tasks`](https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks) | config: splice_site_acceptor; split: default | 4431 (validation) | rna | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_rna_splice_site_donor` | [`genbio-ai/rna-downstream-tasks`](https://huggingface.co/datasets/genbio-ai/rna-downstream-tasks) | config: splice_site_donor; split: default | 4389 (validation) | rna | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_smiles_caption_mol2text` | [`zjunlp/Mol-Instructions`](https://huggingface.co/datasets/zjunlp/Mol-Instructions) | split: default | - | chemistry | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_traitgym_mendelian_dna` | [`bolinas-dna/evals-traitgym_mendelian_v2_harness_255`](https://huggingface.co/datasets/bolinas-dna/evals-traitgym_mendelian_v2_harness_255) | split: default | - | dna | text/structured | classification | exactMatch | no | yes | yes | no | `load_hf_benchmark_tasks` | | `hf_uspto_reaction_prediction` | [`bing-yan/USPTO`](https://huggingface.co/datasets/bing-yan/USPTO) | split: test | - | chemistry | text/structured | qa | openText | no | yes | yes | no | `load_hf_benchmark_tasks` | ### Deprecated aliases | Alias | Source | Config | Split | Count | Domain | Task type | Answer/scorer | Gated | Network | Offline cache | Multimodal | | --- | --- | --- | --- | ---: | --- | --- | --- | --- | --- | --- | --- | | `hf_blue_benchmark` -> `hf_blurb` | [`EMBO/BLURB`](https://huggingface.co/datasets/EMBO/BLURB) | | test | - | biomedical | classification | exactMatch | no | yes | yes | no | | `hf_chinese_medbench` -> `hf_cmb` | [`FreedomIntelligence/CMB`](https://huggingface.co/datasets/FreedomIntelligence/CMB) | CMB-Exam | test | - | medical | mcq | multipleChoice | no | yes | yes | no | | `hf_lavita_medmcqa` -> `medmcqa` | core benchmark alias | | default | loader default | medical | mcq | multipleChoice | no | source-dependent | yes | no | | `hf_lavita_usmle_step1` -> `medqa` | core benchmark alias | | default | loader default | medical | mcq | multipleChoice | no | source-dependent | yes | no | | `hf_lavita_usmle_step2` -> `medqa` | core benchmark alias | | default | loader default | medical | mcq | multipleChoice | no | source-dependent | yes | no | | `hf_lavita_usmle_step3` -> `medqa` | core benchmark alias | | default | loader default | medical | mcq | multipleChoice | no | source-dependent | yes | no | | `hf_mednli_augmented` -> `hf_mednli` | [`araag2/MedNLI`](https://huggingface.co/datasets/araag2/MedNLI) | processed | test | - | clinical | classification | exactMatch | no | yes | yes | no | | `hf_openddi` -> `hf_ddi_corpus_2013` | [`OpenMed/DDI-Corpus-Processed`](https://huggingface.co/datasets/OpenMed/DDI-Corpus-Processed) | | test | - | biomedical | classification | exactMatch | no | yes | yes | no | | `hf_pubmed_20k_rct` -> `hf_pubmed_200k_rct` | [`pietrolesci/pubmed-200k-rct`](https://huggingface.co/datasets/pietrolesci/pubmed-200k-rct) | | default | - | biomedical | classification | exactMatch | no | yes | yes | no | | `hf_pubmed_rct20k` -> `hf_pubmed_200k_rct` | [`pietrolesci/pubmed-200k-rct`](https://huggingface.co/datasets/pietrolesci/pubmed-200k-rct) | | default | - | biomedical | classification | exactMatch | no | yes | yes | no | | `hf_usmle_step_series` -> `medqa` | core benchmark alias | | default | loader default | medical | mcq | multipleChoice | no | source-dependent | yes | no |