Command-Line Reference¶
The basic form of the opencompass command is:
opencompass [config] [options]
This page reflects the current implementation in opencompass/cli/main.py. If the code changes, treat the output of the following command as authoritative:
opencompass --help
Configuration Entry Points and Lookup¶
Argument |
Default |
Description |
|---|---|---|
|
None |
Optional positional argument specifying the path to a Python configuration file. |
|
— |
Display help information. |
|
None |
Find and load one or more model configurations by name. |
|
None |
Find and load one or more dataset configurations by name. |
|
|
Select a result-summary configuration in shorthand configuration mode. Use |
|
|
Specify a custom configuration root. OpenCompass searches its |
When config is supplied, OpenCompass reads that file first. Shorthand construction arguments such as --models, --datasets, --summarizer, --hf-*, and --custom-dataset-* do not replace its model or dataset configurations. Without a configuration file, use one of these entry points:
load existing configurations with
--modelsand--datasets;construct a Hugging Face model with
--hf-pathand select datasets with--datasets;select a model with
--modelsor--hf-pathand construct a custom dataset with--custom-dataset-path.
Execution Stages and Working Directory¶
Argument |
Default |
Description |
|---|---|---|
|
|
Select the execution stage: |
|
None |
Reuse an experiment directory for the specified timestamp. If the timestamp is omitted, use the last directory by name in the working directory. |
|
|
Set the experiment output root. Artifacts are stored in its timestamped subdirectory. |
|
|
Enable debug mode. The Runner executes tasks sequentially and displays logs directly, which is useful for diagnosing a first run. |
|
|
Parse the configuration and partition tasks without starting inference or evaluation. This also enables debug-level logging. |
|
None |
Attempt to convert supported local Hugging Face model configurations to vLLM or LMDeploy for one-stop deployment and evaluation. Unsupported model types remain unchanged and produce a warning. |
|
|
Print the loaded and processed experiment configuration. |
|
|
Enable Lark bot task notifications. The configuration must also provide |
Task Partitioning and Runner¶
Argument |
Default |
Description |
|---|---|---|
|
|
Set the maximum concurrent task count for a default Runner supplied by the entry point and set |
|
|
Set the maximum number of concurrent tasks per GPU for an automatically generated |
|
|
Force |
|
|
Force the Alibaba Cloud PAI-DLC Runner and replace existing |
Inference, Evaluation, and Analysis Outputs¶
Argument |
Default |
Description |
|---|---|---|
|
|
Save per-sample evaluation details. Use |
|
|
Pass the response-length statistics flag to inference tasks. Support depends on the Inferencer. |
|
None |
Export constructed messages without requesting the model. Currently supported only by |
|
|
Instruct evaluation tasks to calculate and save the answer extraction rate. |
|
|
Analyze repeated predictions during summarization and write a repeated-output analysis file. |
|
|
In shorthand CLI configuration mode, set |
Quickly Constructing a Hugging Face Model¶
The following arguments apply only when neither config nor --models supplies a model and --hf-path is used to construct one:
Argument |
Default |
Description |
|---|---|---|
|
|
Select the base-model or chat-model wrapper. |
|
None |
Specify a Hugging Face model path or repository ID. |
|
|
Pass model-loading arguments. |
|
Same as model path |
Specify a tokenizer path or repository ID. |
|
|
Pass tokenizer-loading arguments. |
|
None |
Specify a PEFT weights path. |
|
|
Pass PEFT-loading arguments. |
|
|
Pass generation arguments. |
|
None |
Set the maximum sequence length supported by the model. |
|
|
Set the maximum output length. |
|
|
Set the minimum output length. |
|
|
Set the inference batch size. |
|
|
Set the number of GPUs used by a Hugging Face model task. |
|
None |
Specify the padding token ID. |
|
Empty list |
Specify one or more stop words. |
|
— |
Deprecated. The current version raises an error when this argument is supplied; use |
Quickly Constructing a Custom Dataset¶
The following arguments generate a dataset configuration from a local file:
Argument |
Default |
Description |
|---|---|---|
|
None |
Specify a custom dataset file. Required when neither |
|
None |
Specify the metadata file for the custom dataset. |
|
Auto-detected |
Select multiple-choice or question-answer data. |
|
Auto-detected |
Select generative or PPL inference. |
Result Persistence Arguments¶
Argument |
Default |
Description |
|---|---|---|
|
None |
Specify a shared result directory. When this argument or |
|
|
Read existing results from the shared directory before execution, write them into the current experiment’s |
|
|
Allow existing files to be overwritten when saving results to the shared directory. |