Add a dataset¶
Although OpenCompass has already included most commonly used datasets, users need to follow the steps below to support a new dataset if wanted:
Add a dataset script
mydataset.pyto theopencompass/datasetsfolder. This script should include:The dataset and its loading method. Define a
MyDatasetclass that implements the data loading methodloadas a static method. This method should return data of typedatasets.Datasetordatasets.DatasetDict. We use the Hugging Face dataset as the unified interface for datasets to avoid introducing additional logic. Ifloadreturns aDataset, OpenCompass will use it as both internaltrainandtestsplits. If it returns aDatasetDict, you can specify the actual splits withtrain_splitandtest_splitinreader_cfg. Here’s an example:
from typing import Union import datasets from opencompass.registry import LOAD_DATASET from .base import BaseDataset @LOAD_DATASET.register_module() class MyDataset(BaseDataset): @staticmethod def load(**kwargs) -> Union[datasets.Dataset, datasets.DatasetDict]: pass
(Optional) If the existing evaluators in OpenCompass do not meet your needs, you can implement and register a custom Evaluator. See Adding Postprocessors, Evaluators, and Summarizers for details.
(Optional) If the existing postprocessors in OpenCompass do not meet your needs, you need to define the
mydataset_postprocessmethod. This method takes an input string and returns the corresponding postprocessed result string. If you want to reuse the postprocessor by a registry name, register it toTEXT_POSTPROCESSORS. Here’s an example:
from opencompass.registry import TEXT_POSTPROCESSORS @TEXT_POSTPROCESSORS.register_module('mydataset') def mydataset_postprocess(text: str) -> str: pass
After adding the dataset script, make sure the related classes and functions can be imported by the config file. If you want to use
from opencompass.datasets import ..., import the new module inopencompass/datasets/__init__.py; alternatively, import directly from the concrete module in the config file, for examplefrom opencompass.datasets.mydataset import MyDataset.After defining the dataset loading, data postprocessing, and evaluator methods, you need to add the following configurations to the configuration file:
from opencompass.datasets import MyDataset, MyDatasetEvaluator, mydataset_postprocess mydataset_eval_cfg = dict( evaluator=dict(type=MyDatasetEvaluator), pred_postprocessor=dict(type=mydataset_postprocess)) mydataset_datasets = [ dict( type=MyDataset, ..., reader_cfg=..., infer_cfg=..., eval_cfg=mydataset_eval_cfg) ]
To make your dataset easier for other users to access, specify the dataset path in the configuration file. The
pathfield can be a local path or a logical dataset name. A logical dataset name is resolved through the mapping inopencompass/utils/datasets_info.py. Here’s an example:
mmlu_datasets = [ dict( ..., path='opencompass/mmlu', ..., ) ]
Next, you need to create a dictionary key in
opencompass/utils/datasets_info.pywith the same name as the one you provided above. If you have already hosted the dataset on Hugging Face or ModelScope, please add a dictionary key to theDATASETS_MAPPINGdictionary and fill in the Hugging Face or ModelScope dataset address in thehf_idorms_idkey, respectively. You can also specify a defaultlocaladdress. Here’s an example:
"opencompass/mmlu": { "ms_id": "opencompass/mmlu", "hf_id": "opencompass/mmlu", "local": "./data/mmlu/", }
If you wish for the provided dataset to be accessible through the OpenCompass OSS repository when used by others, you need to submit the dataset files in the Pull Request phase. We will then transfer the dataset to the OSS on your behalf and create a new dictionary key in
DATASETS_URL.To keep data sources selectable, implement the
loadmethod inmydataset.pyaccording to the path type you provide. Usually, callget_data_path(path)first to resolve the path: whenDATASET_SOURCE=ModelScope, it usesms_id; whenDATASET_SOURCE=HF, it useshf_id; whenDATASET_SOURCEis not set, the current implementation first uses the local path from thelocalfield and combines it withCOMPASS_DATA_CACHEto look for cached data. It only tries to download throughDATASETS_URLwhen the local path does not exist. If different data sources return different data formats, adapt them inload. Here’s an example fromopencompass/datasets/cmmlu.py:
def load(path: str, name: str, **kwargs): ... if environ.get('DATASET_SOURCE') == 'ModelScope': ... else: ... return dataset
After completing the dataset script and config file, you need to register the information of your new dataset in the file
dataset-index.ymlat the main directory, so that it can be added to the dataset statistics list on the OpenCompass website.The keys that need to be filled in include
name: the name of your dataset,category: the category of your dataset,paper: the URL of the paper or project,configpath: the path to the dataset config file, andconfigpath_llmjudge: the path to the LLM Judge config file. If no LLM Judge config is available, setconfigpath_llmjudgeto an empty string. Here’s an example:
- mydataset: name: MyDataset category: Understanding paper: https://arxiv.org/pdf/xxxxxxx configpath: opencompass/configs/datasets/MyDataset configpath_llmjudge: ''
Detailed dataset configuration files and other required configuration files can be referred to in the Configuration Files tutorial. For guides on launching tasks, please refer to the Quick Start tutorial.