Dataset Selection and Configuration¶
An OpenCompass dataset configuration defines data loading, model input construction, and scoring rules together. The same dataset may have multiple configuration variants, and different evaluation protocols may produce different results.
Selecting a Configuration¶
python tools/list_configs.py mmlu gsm8k # Find configurations related to MMLU and GSM8K
Dataset configuration files are usually located under opencompass/configs/datasets/<dataset>/. Their filenames commonly include identifiers such as gen, ppl, rawprompt, the number of few-shot examples, and a hash to distinguish evaluation protocols.
Before selecting a configuration, verify the following:
data source, version, split, and sample range;
input fields, answer field, and any multimodal input;
prompt type, number of few-shot examples, and inference method;
Evaluator and post-processing rules, including any dependency on a Judge model or an external evaluation service.
Dataset Configuration Structure¶
A dataset configuration consists of data-loading arguments and three sections: reader_cfg, infer_cfg, and eval_cfg.
datasets = [
dict(
type=MyDataset,
abbr='my-dataset',
path='data/or/hub-id',
reader_cfg=reader_cfg,
infer_cfg=infer_cfg,
eval_cfg=eval_cfg,
)
]
type: the dataset class registered with OpenCompass. It loads the source data into a Hugging FaceDatasetorDatasetDict.abbr: the short name used in task directories and summary results. Use distinguishable abbreviations for different evaluation configurations of the same source data.path: a dataset path or repository identifier. Depending on the dataset class, arguments such asname,split, ortaskmay also be accepted to select a subset.reader_cfg: specifies the fields, data splits, and sample ranges used in evaluation.infer_cfg: specifies prompt construction, few-shot example retrieval, and the inference method.eval_cfg: specifies prediction post-processing and scoring rules.
reader_cfg: Fields and Data Splits¶
The basic form of reader_cfg is as follows:
reader_cfg = dict(
input_columns=['question'],
output_column='answer',
train_split='train',
test_split='test',
train_range=None,
test_range='[:100]',
)
The fields have the following meanings:
input_columns: fields used to construct the model input, such as the question, options, or context.output_column: the field containing the reference answer. Set it toNonefor tasks that do not require reference answers.train_splitandtest_split: respectively specify the split from which the Retriever selects few-shot examples and the split on which inference and scoring are performed. Their defaults aretrainandtest. These fields may be omitted when the data has only one split.train_rangeandtest_range: limit the samples selected from the corresponding splits.Noneuses all samples; an integer selects a fixed number after shuffling; a float between0and1selects that proportion; and a slice string such as'[:100]'or'[100:200]'selects an interval in the original order. These fields may be omitted for full evaluation.
After completing the configuration, verify that input_columns, output_column, and every placeholder in the prompt exist in the loaded data. To quickly test the first several samples, use test_range='[:N]'.
infer_cfg: Prompt, Retrieval, and Inference¶
The most common structure for generative evaluation is:
from opencompass.openicl.icl_inferencer import GenInferencer
from opencompass.openicl.icl_raw_prompt_template import RawPromptTemplate
from opencompass.openicl.icl_retriever import ZeroRetriever
infer_cfg = dict(
prompt_template=dict(
type=RawPromptTemplate,
messages=[
dict(role='user', content='{question}\nPlease provide the answer.'),
],
),
retriever=dict(type=ZeroRetriever),
inferencer=dict(type=GenInferencer),
)
infer_cfg commonly contains the following fields:
prompt_template: converts each dataset sample into model input through a prompt template.retriever: determines which few-shot examples are selected from the training split.ZeroRetrieverretrieves no examples;FixKRetriever,RandomRetriever, and other Retrievers select examples according to their respective strategies.inferencer: determines the inference method.GenInferencerasks the model to generate an answer directly and accepts inference arguments such asmax_out_lenandstopping_criteria. Amax_out_lenexplicitly set here takes precedence over the default in the model configuration.
For prompt placeholders, conversation messages, and few-shot insertion, see Prompt Templates. After changing a template, use Prompt Preview and Debugging to inspect the final input.
A common PPL evaluation configuration is:
from opencompass.openicl.icl_inferencer import PPLInferencer
from opencompass.openicl.icl_prompt_template import PromptTemplate
from opencompass.openicl.icl_retriever import ZeroRetriever
infer_cfg = dict(
prompt_template=dict(
type=PromptTemplate,
template={
'yes': '{question} yes',
'no': '{question} no',
},
),
retriever=dict(type=ZeroRetriever),
inferencer=dict(type=PPLInferencer),
)
The keys of the candidate templates must cover the label values in output_column. Alternatively, candidate labels may be specified explicitly through the labels argument of PPLInferencer. PPL evaluation requires a model backend capable of computing log-likelihoods for its input; not every API model provides this capability.
eval_cfg: Post-processing and Scoring¶
When model outputs cannot be compared directly with reference answers, each side can be post-processed before scoring. For example:
from opencompass.datasets import (Gsm8kEvaluator,
gsm8k_dataset_postprocess,
gsm8k_postprocess)
eval_cfg = dict(
evaluator=dict(type=Gsm8kEvaluator),
pred_postprocessor=dict(type=gsm8k_postprocess),
dataset_postprocessor=dict(type=gsm8k_dataset_postprocess),
)
The fields serve the following purposes:
evaluator: the scorer configuration.typeselects the Evaluator, and all remaining fields are passed as initialization arguments. Common scoring methods include accuracy, exact match, mathematical answer verification, code execution, and model-based judging.pred_postprocessor: processes model predictions before scoring, for example by extracting an option letter, a number, or an answer enclosed in a particular tag.dataset_postprocessor: processes reference answers fromoutput_columnbefore scoring so that their format matches the processed predictions.pred_role: extracts content for a specified role from the output of a local chat model. Use it only when the model defines a correspondingmeta_template.
An Evaluator type is required by the basic scoring workflow; all other fields are optional.
For data caching and offline behavior, see Data Sources, Caching, and Offline Operation. For the complete dataset extension procedure, see Adding a Dataset.
Dataset Statistics¶
The following table is generated from dataset-index.yml in the repository root. It lists the datasets registered with OpenCompass, their categories, resource links, and recommended configurations, and supports fuzzy search.
Name |
Category |
Paper or Repository |
Recommended Config |
Recommended Config (LLM Judge) |
|---|---|---|---|---|
IFEval |
Instruction Following |
|||
Inverse IFEval |
Instruction Following |
|||
NPHardEval |
Reasoning |
|||
PMMEval |
Language |
|||
PI-LLM |
Memory |
|||
TheoremQA |
Reasoning |
|||
AGIEval |
Examination |
|||
BABILong |
Long Context |
|||
AA-LCR |
Long Context |
|||
BigCodeBench |
Code |
|||
CaLM |
Reasoning |
|||
InfiniteBench (∞Bench) |
Long Context |
|||
KOR-Bench |
Reasoning |
|||
LawBench |
Knowledge / Law |
|||
L-Eval |
Long Context |
|||
LiveCodeBench |
Code |
|||
LiveCodeBench Pro |
Code |
|||
LiveMathBench |
Math |
|||
LiveReasonBench |
Reasoning |
|||
LongBench |
Long Context |
|||
LV-Eval |
Long Context |
|||
Mastermath2024v1 |
Math |
|||
matbench |
Science / Material |
|||
MedBench |
Knowledge / Medicine |
|||
MedCalc_Bench |
Knowledge / Medicine |
|||
MedQA |
Knowledge / Medicine |
|||
MedXpertQA |
Knowledge / Medicine |
|||
ClinicBench |
Knowledge / Medicine |
|||
ScienceQA |
Knowledge / Medicine |
|||
PubMedQA |
Knowledge / Medicine |
|||
MuSR |
Reasoning |
|||
NeedleBench V1 (Deprecated) |
Long Context |
|||
NeedleBench V2 |
Long Context |
|||
RULER |
Long Context |
|||
AlignBench |
Subjective / Alignment |
|||
AlpacaEval |
Subjective / Instruction Following |
|||
Arena-Hard |
Subjective / Chatbot |
|||
ELBench |
Subjective / Education |
|||
FLAMES |
Subjective / Alignment |
|||
FOFO |
Subjective / Format Following |
|||
FollowBench |
Subjective / Instruction Following |
|||
HelloBench |
Subjective / Long Context |
|||
JudgerBench |
Subjective / Long Context |
|||
MT-Bench-101 |
Subjective / Multi-Round |
|||
WildBench |
Subjective / Real Task |
|||
T-Eval |
Tool Utilization |
|||
BuySideFinBench |
Knowledge / Finance |
|||
FinanceIQ |
Knowledge / Finance |
|||
GAOKAOBench |
Examination |
|||
LCBench |
Code |
|||
ArabicMMLU |
Language |
|||
OpenFinData |
Knowledge / Finance |
|||
QuALITY |
Long Context |
|||
Adversarial GLUE |
Safety |
link(TBD) / link(TBD) / link(TBD) / link(TBD) / link(TBD) / link(TBD) |
||
CLUE / AFQMC |
Language |
|||
AIME2024 |
Examination |
|||
Adversarial NLI |
Reasoning |
|||
Anthropics Evals |
Safety |
|||
APPS |
Code |
|||
ARC |
Reasoning |
|||
ARC Prize |
ARC-AGI |
|||
ArxivRollBench |
Reasoning / Robustness |
|||
SuperGLUE / AX |
Reasoning |
|||
BIG-Bench Hard |
Reasoning |
|||
BIG-Bench Extra Hard |
Reasoning |
|||
SuperGLUE / BoolQ |
Knowledge |
|||
CLUE / C3 (C³) |
Understanding |
|||
CARDBiomedBench |
Knowledge / Medicine |
|||
SuperGLUE / CB |
Reasoning |
|||
C-EVAL |
Examination |
|||
CHARM |
Reasoning |
|||
ChemBench |
Knowledge / Chemistry |
|||
FewCLUE / CHID |
Language |
|||
Chinese SimpleQA |
Knowledge |
|||
CIBench |
Code |
|||
CivilComments |
Safety |
|||
Cloze Test-max/min |
Code |
|||
FewCLUE / CLUEWSC |
Language / WSC |
|||
CMB |
Knowledge / Medicine |
|||
CMMLU |
Understanding |
|||
CLUE / CMNLI |
Reasoning |
|||
cmo_fib |
Examination |
|||
CLUE / CMRC |
Understanding |
|||
CommonSenseQA |
Knowledge |
|||
CommonSenseQA-CN |
Knowledge |
|||
SuperGLUE / COPA |
Reasoning |
|||
CrowsPairs |
Safety |
|||
CrowsPairs-CN |
Safety |
|||
CVALUES |
Safety |
|||
CLUE / DRCD |
Understanding |
|||
DROP (DROP Simple Eval) |
Understanding |
|||
DS-1000 |
Code |
|||
FewCLUE / EPRSTMT |
Understanding |
|||
Flores |
Language |
|||
Game24 |
Math |
|||
Government Report Dataset |
Long Context |
|||
GPQA |
Knowledge |
|||
GSM8K |
Math |
|||
GSM-Hard |
Math |
|||
HLE(Humanity’s Last Exam) |
Reasoning |
|||
HellaSwag |
Reasoning |
|||
HumanEval |
Code |
|||
HumanEval-CN |
Code |
|||
Multi-HumanEval |
Code |
|||
HumanEval+ |
Code |
|||
HumanEval-X |
Code |
|||
HumanEval Pro |
Code |
|||
Hungarian_Math |
Math |
|||
IWSLT2017 |
Language |
|||
JigsawMultilingual |
Safety |
|||
LAMBADA |
Understanding |
|||
LCSTS |
Understanding |
|||
LiveStemBench |
||||
LLM Compression |
Bits Per Character (BPC) |
|||
MATH |
Math |
|||
MATH500 |
Math |
|||
MATH 401 |
Math |
|||
MathBench |
Math |
|||
MBPP |
Code |
|||
MBPP-CN |
Code |
|||
MBPP-PLUS |
Code |
|||
MBPP Pro |
Code |
|||
MGSM |
Language / Math |
|||
MMLU |
Understanding |
|||
SciEval |
Understanding |
|||
MMLU-CF |
Understanding |
|||
MMLU-Pro |
Understanding |
|||
MMMLU |
Language / Understanding |
|||
SuperGLUE / MultiRC |
Understanding |
|||
MultiPL-E |
Code |
|||
NarrativeQA |
Understanding |
|||
NaturalQuestions |
Knowledge |
|||
NaturalQuestions-CN |
Knowledge |
|||
OpenBookQA |
Knowledge |
|||
OlymMATH |
Math |
|||
OpenBookQA |
Knowledge / Physics |
|||
ProteinLMBench |
Knowledge / Biology (Protein) |
|||
py150 |
Code |
|||
Qasper |
Long Context |
|||
Qasper-Cut |
Long Context |
|||
RACE |
Examination |
|||
R-Bench |
Reasoning |
|||
RealToxicPrompts |
Safety |
|||
SuperGLUE / ReCoRD |
Understanding |
|||
SuperGLUE / RTE |
Reasoning |
|||
CLUE / OCNLI |
Reasoning |
|||
FewCLUE / OCNLI-FC |
Reasoning |
|||
RoleBench |
Role Play |
|||
S3Eval |
Long Context |
|||
SciBench |
Reasoning |
|||
SciCode |
Code |
|||
SeedBench |
Knowledge |
|||
SimpleQA |
Knowledge |
|||
SocialIQA |
Reasoning |
|||
SQuAD2.0 |
Understanding |
|||
StoryCloze |
Reasoning |
|||
StrategyQA |
Reasoning |
|||
SummEdits |
Language |
|||
SummScreen |
Understanding |
|||
SVAMP |
Math |
|||
TabMWP |
Math / Table |
|||
TACO |
Code |
|||
FewCLUE / TNEWS |
Understanding |
|||
FewCLUE / BUSTM |
Reasoning |
|||
FewCLUE / CSL |
Understanding |
|||
FewCLUE / OCNLI-FC |
Reasoning |
|||
TriviaQA |
Knowledge |
|||
TriviaQA-RC |
Knowledge / Understanding |
|||
TruthfulQA |
Safety |
|||
TyDi-QA |
Language |
|||
SuperGLUE / WiC |
Language |
|||
SuperGLUE / WSC |
Language / WSC |
|||
WinoGrande |
Language / WSC |
|||
XCOPA |
Language |
|||
Xiezhi |
Knowledge |
|||
XLSum |
Understanding |
|||
Xsum |
Understanding |
|||
GLUE / CoLA |
Understanding |
|||
GLUE / MPRC |
Understanding |
|||
GLUE / QQP |
Understanding |
|||
Omni-MATH |
Math |
|||
WikiBench |
Knowledge |
|||
SuperGPQA |
Knowledge |
|||
ClimaQA |
Science |
|||
PHYSICS |
Science |
|||
SmolInstruct |
Science /Chemistry |
|||
SciKnowEval |
Science |
|||
InternSandbox |
Reasoning/Code/Agent |
|||
nejmaibench |
Science /Medicine |
|||
Medbullets |
Science /Medicine |
|||
medmcqa |
Science /Medicine |
|||
PHYBench |
Science /Physics |
|||
BeyondAIME |
Math |
|||
EESE |
Science |
|||
CL-bench |
Long Context |
|||
CMPhysBench |
Science /Physics |
|||
Earth-Silver |
Science |
|||
HealthBench |
Knowledge / Medicine |
|||
IFBench |
Instruction Following |
|||
Mol-Instructions (Chem) |
Knowledge / Chemistry |
|||
OlympiadBench |
Math |
|||
ProcessBench |
Math |
|||
SciReasoner |
Science |
|||
SciReasoner1.5 |
Science |
|||
AdvancedIF |
Instruction Following |
|||
AIME2025 |
Math |
|||
AIME2026 |
Math |
|||
ATLAS |
Science |
|||
CodeCompass |
Code |
|||
JudgeBench |
Subjective / Alignment |
|||
JudgerBenchV2 |
Subjective / Alignment |
|||
RewardBench |
Subjective / Alignment |
|||
RMB |
Subjective / Alignment |
|||
MolecularIQ |
Science /Chemistry |
|||
MP-20 |
Science / Material |
|||
MRCR |
Long Context |
|||
OJBench |
Code |
|||
OpenSWI |
Science |
|||
PerspectiveGap |
Reasoning/Code/Agent |
|||
PromptBench |
Safety |
|||
S2-TOMG-Bench |
Science /Chemistry |
|||
SRBench |
Science |
|||
WikiText |
Language |
|||
Winograd Schema Challenge |
Language / WSC |
|||
WritingBench |
Subjective / Writing |
|||
Biology Instructions |
Knowledge / Biology |
|||
DINGO |
Instruction Following |
|||
HMMT2026 |
Math |
|||
ZebraLogic |
Reasoning |