Multimodal Evaluation Overview¶
OpenCompass reuses multimodal datasets and official evaluators from VLMEvalKit through a bridge. Dataset construction and scoring use the official VLMEvalKit implementation, while model inference, task scheduling, and result summarization use standard OpenCompass workflows. In addition to text configuration, multimodal evaluation involves image download locations, media blocks in messages, and image-input support in the model backend.
Integration: Responsibilities of Three Components¶
The bridge consists of VLMEvalKitDataset on the data side, the standard inference pipeline, and VLMEvalKitEvaluator on the scoring side:
Data loading:
VLMEvalKitDataset.loadcalls VLMEvalKitbuild_dataset(dataset_name)to construct the official dataset object. On first use, TSV data and images are downloaded under theLMUDatacache according to official rules. Officialbuild_prompt()then builds each input and the bridge converts it into structured OpenCompass messages: text blocks{'type': 'text', ...}and image blocks{'type': 'image', 'image_url': <local path or URL>}. Every sample also recordssample_id(<dataset name>:<index>), the original row JSON, and reference answer.Inference: the standard OpenCompass workflow is used. Dataset
infer_cfgpasses messages in thepromptcolumn directly to the model using RawPromptTemplateexpand_column, thenGenInferencergenerates normally. The backend must support image-content blocks.OpenAI/OpenAISDKcurrently do: local image paths are converted to base64 data URLs automatically, whileimage_formatandimage_min_edgecontrol re-encoding and minimum resolution.Scoring:
VLMEvalKitEvaluatoraligns predictions to the official table bysample_id, exports xlsx, and calls officialdataset.evaluate(). It flattens returned aggregate metrics, converts them to percentages, and returns them to standard OpenCompass results and summarization.
The bridge has explicit boundaries and raises during construction otherwise:
Only IMAGE-modality datasets are supported; video is unsupported.
Only single-turn datasets are supported; official datasets marked as requiring multi-turn inference (
TYPEisMT) are unsupported.A unique, nonempty
indexcolumn is required.
Installation¶
pip install "opencompass[vlm]"
The vlm extra installs libraries needed on the multimodal model side, including litellm and google-genai. VLMEvalKit itself must be installed separately according to its official repository, so that import vlmeval works. Python 3.10+ is required; the bridge checks both requirements at startup.
Running the Two Integrated Datasets¶
Configurations are currently provided for MMBench (DEV_EN) and MMMU-Pro (10c), together with complete examples:
Dataset configurations: opencompass/configs/datasets/MMBench/MMBench_DEV_EN_vlmevalkit_gen.py and opencompass/configs/datasets/MMMU_Pro/MMMU_Pro_10c_vlmevalkit_gen.py.
End-to-end examples: examples/eval_mmbench_vlmevalkit.py and examples/eval_mmmu_pro_vlmevalkit.py.
For MMBench, first replace the example’s tested-model endpoint and official-scoring LLM endpoint with usable services, then run:
export OPENAI_API_KEY=sk-xxx
opencompass examples/eval_mmbench_vlmevalkit.py
The https://example.com URLs in the example are placeholders and cannot be used for formal evaluation as-is.
The complete workflow has four steps:
Load:
build_datasetdownloads/reads official TSV and images underLMUData, then generates messages with image blocks.Infer:
expand_columnpasses messages to the model. The example usesOpenAISDKagainst an OpenAI-compatible multimodal endpoint, automatically converting local images to base64.Score: predictions are aligned to the official table, exported to xlsx, and given to official
evaluate(). MMBench official scoring internally uses an LLM to extract choices, so the example setsmodel,api_base,nproc,retry,timeout, and related fields undereval_cfg.evaluator.eval_kwargs; they are forwarded unchanged to official scoring.Summarize: metrics enter standard result files and summary tables.
For a small trial, limit samples through an environment variable:
MMBENCH_SAMPLE_LIMIT=20 opencompass examples/eval_mmbench_vlmevalkit.py
To use your own OpenAI-compatible multimodal service, change path and openai_api_base in the example model configuration and retain image arguments such as image_format. If the dataset’s official scoring requires additional LLM calls, such as MMBench choice extraction, also update scoring arguments such as model and api_base under eval_cfg.evaluator.eval_kwargs. A language-model configuration cannot be substituted by changing only type; the model must actually accept image input.
Data Cache and Environment Variable¶
Environment variable |
Purpose |
Default |
|---|---|---|
|
VLMEvalKit data-cache root; TSV and images are downloaded here |
|
LMUData is VLMEvalKit’s own data-directory convention. The dataset configuration reads it as data_root; during dataset construction and official scoring, the bridge temporarily points LMUData to this directory. Relative paths become absolute and are created automatically, ensuring that download, image reads, and scoring share the same data. On shared storage or in a container, explicitly set and mount it:
export LMUData=/shared/cache/vlmevalkit
Reading Results¶
Assume the configured work_dir is outputs/mmbench_vlmevalkit and the model abbreviation is kimi-k2.6-chat-completions. OpenCompass appends a timestamp directory under work_dir; below, <exp_dir> means the actual experiment directory, for example outputs/mmbench_vlmevalkit/20260916_120000:
Predictions:
<exp_dir>/predictions/<model abbr>/MMBench_DEV_EN.json; each record contains original input messages including image references and the model reply.Scoring artifacts:
<exp_dir>/results/<model abbr>/MMBench_DEV_EN.jsonis the standard OpenCompass metric file. The siblingMMBench_DEV_EN/directory always contains the following bridge artifacts, while officialevaluate()may generate additional files such as MMBench_acc.csv:MMBench_DEV_EN.xlsx: complete prediction table aligned to official data, used directly by official scoring.vlmevalkit_evaluation.json: snapshot of official scoring arguments including dataset name, data directory, andeval_kwargs, for reproduction.vlmevalkit_metrics.json: flattened official aggregate metrics and the primary metric.
Summary: CSV aggregate table under
<exp_dir>/summary/.
Metric names follow flattened official VLMEvalKit output, including group columns and Overall, and values are normalized to percentages. Dataset official logic determines the primary metric. Any empty sample prediction makes scoring fail because official scoring requires a complete prediction sequence.