Subjective Evaluation Guidance¶
Introduction¶
Subjective evaluation aims to assess the model’s performance in tasks that align with human preferences. The key criterion for this evaluation is human preference, but it comes with a high cost of annotation.
To explore the model’s subjective capabilities, we employ JudgeLLM as a substitute for human assessors (LLM-as-a-Judge).
A popular evaluation method involves
Compare Mode: comparing model responses pairwise to calculate their win rate
Score Mode: another method involves calculate scores with single model response (Chatbot Arena).
Based on these methods, OpenCompass supports JudgeLLM-based subjective evaluation. Any model supported by the OpenCompass repository can be used directly as a JudgeLLM, and support for additional specialized JudgeLLMs is also planned.
Currently Supported Subjective Evaluation Datasets¶
AlignBench Chinese Scoring Dataset
MTBench English Scoring Dataset, two-turn dialogue
MTBench101 English Scoring Dataset, multi-turn dialogue
AlpacaEvalv2 English Compare Dataset
ArenaHard English Compare Dataset, mainly focused on coding
Fofo English Scoring Dataset
Wildbench English Score and Compare Dataset
CompassArena Chinese Compare Dataset
CompassArena-SubjectiveBench single-turn and multi-turn Compare Dataset with Bradley-Terry summarization
CompassBench Chinese and English Compare Dataset
ELBench education-focused evaluation Dataset with LLM-as-a-Judge subjective subsets
FLAMES Chinese Alignment Scoring Dataset
FollowBench Chinese and English Instruction Following Scoring Dataset
HelloBench Long Text Generation Scoring Dataset
WritingBench Writing Scoring Dataset
Initiating Subjective Evaluation¶
Similar to existing objective evaluation methods, you can configure related settings in examples/eval_subjective.py.
Basic Parameters: Specifying models, datasets, and judgemodels¶
Similar to objective evaluation, import the models and datasets that need to be evaluated, for example:
with read_base():
from .datasets.subjective.alignbench.alignbench_judgeby_critiquellm_rawprompt import alignbench_datasets
from .datasets.subjective.alpaca_eval.alpacav2_judgeby_gpt4 import subjective_datasets as alpacav2
from .models.openai.gpt_6_astra import models
Specifying Other Parameters¶
In addition to the basic parameters, you can also modify the infer and eval fields in the config to set a more appropriate partitioning method. The currently supported partitioning methods mainly include three types: NaivePartitioner, SizePartitioner, and NumberWorkPartitioner. You can also specify your own workdir to save related files.
Subjective Evaluation with Custom Dataset¶
The specific process includes:
Data preparation
Model response generation
Evaluate the response with a JudgeLLM
Generate JudgeLLM’s response and calculate the metric
Step-1: Data Preparation¶
This step requires preparing the dataset file and implementing your own dataset class under Opencompass/datasets/subjective/, returning the read data in the format of list of dict.
Actually, you can prepare the data in any format you like (csv, json, jsonl, etc.). However, to make it easier to get started, it is recommended to construct the data according to the format of the existing subjective datasets or according to the following json format. We provide mini test-set for Compare Mode and Score Mode as below:
### Compare-Mode Example
[
{
"question": "If I throw a ball vertically into the air, which direction does it initially travel?",
"capability": "Knowledge - common sense",
"others": {
"question": "If I throw a ball vertically into the air, which direction does it initially travel?",
"evaluating_guidance": "",
"reference_answer": "Up"
}
},...]
### Score-Mode Dataset Example
[
{
"question": "Act as an email assistant. Draft an approximately 200-word email asking my advisor whether a research sync can be held at 15:00 next Wednesday.",
"capability": "Email notification",
"others": ""
},
The json must includes the following fields:
‘question’: Question description
‘capability’: The capability dimension of the question.
‘others’: Other needed information.
These three fields are required, and users may add other fields. To customize the prompt for individual questions, add the required information to others and expose the corresponding fields in the Dataset class.
Step-2: Evaluation Configuration(Compare Mode)¶
Taking Alignbench as an example, configs/datasets/subjective/alignbench/alignbench_judgeby_critiquellm_rawprompt.py:
First, you need to set
subjective_reader_cfgto receive the relevant fields returned from the custom Dataset class and specify the output fields when saving files.Then, you need to specify the root path
data_pathof the dataset and the dataset filenamesubjective_all_sets. If there are multiple sub-files, you can add them to this list.Specify
subjective_infer_cfgandsubjective_eval_cfgto configure the corresponding inference and evaluation prompts.Specify additional information such as
modeat the corresponding location. Note that the fields required for different subjective datasets may vary.Define post-processing and score statistics. For example, the
alignbench_postprocessfunction inopencompass/datasets/subjective/alignbench.py.
Step-3: Launch the Evaluation¶
opencompass examples/eval_subjective.py -r
The -r parameter allows the reuse of model inference and GPT-4 evaluation results.
The response of JudgeLLM will be output to output/.../results/timestamp/xxmodel/xxdataset/.json.
The evaluation report will be output to output/.../summary/timestamp/report.csv.
Multi-round Subjective Evaluation in OpenCompass¶
In OpenCompass, we also support subjective multi-turn dialogue evaluation. For instance, the evaluation of MT-Bench can be referred to in configs/datasets/subjective/multiround.
In the multi-turn dialogue evaluation, you need to organize the data format into the following dialogue structure:
"dialogue": [
{
"role": "user",
"content": "Imagine you are participating in a race with a group of people. If you have just overtaken the second person, what's your current position? Where is the person you just overtook?"
},
{
"role": "assistant",
"content": ""
},
{
"role": "user",
"content": "If the \"second person\" is changed to \"last person\" in the above question, what would the answer be?"
},
{
"role": "assistant",
"content": ""
}
],
It’s important to note that due to the different question types in MTBench having different temperature settings, we need to divide the original data files into three different subsets according to the temperature for separate inference. For different subsets, we can set different temperatures. For specific settings, please refer to configs/datasets/subjective/multiround/mtbench_single_judge_diff_temp_new_dialogue.py.