Abstract
Benchmarking video–language models has largely focused on short clips and single-sentence metrics, leaving open whether current systems can generate accurate long-form, paragraph-level descriptions. We introduce CLIP-CC-Bench, an evaluation suite for long-form video description built from 5 hours of movie content segmented into 90-second clips, each paired with an expert-written paragraph-style reference. The suite employs an ensemble of five state-of-the-art LLM-based embedding models to mitigate single-model bias, applying two complementary methodologies — coarse-grained and fine-grained semantic matching — to compare model-generated descriptions against the references. Using this framework we evaluate 17 state-of-the-art video–language models and report both their Borda-aggregated rankings and average scores, and we quantify the protocol's internal reliability through inter-judge agreement and bootstrap ranking stability.
Approach
CLIP-CC-Bench scores a candidate paragraph against an expert reference using an ensemble of five long-context embedding judges. Each judge combines a coarse-grained, paragraph-level cosine similarity with a fine-grained, sentence-level F1 (greedy best-match), merged into a harmonic mean (HM-CF). Per-judge rankings are aggregated with Borda count to produce the leaderboard, and the protocol's reliability is assessed via inter-judge agreement and bootstrap ranking stability.
Results
| Overall Rank | Model | Params | Borda | Mean HM-CF | Mean Coarse | Mean Fine-F1 |
|---|---|---|---|---|---|---|
| 1 | VideoLLaMA3 | 7B | 80 | 0.67 | 0.73 | 0.63 |
| 2 | mPLUG-Owl3 | 7B | 75 | 0.66 | 0.71 | 0.62 |
| 3 | LLaVA-OneVision | 7B | 67 | 0.64 | 0.69 | 0.60 |
| 4 | ViLAMP | 7B | 67 | 0.64 | 0.69 | 0.60 |
| 5 | LongVU | 7B | 61 | 0.63 | 0.68 | 0.59 |
| 6 | Qwen2.5-72B | 72B | 55 | 0.62 | 0.67 | 0.58 |
| 7 | Qwen2.5-32B | 32B | 48 | 0.61 | 0.65 | 0.57 |
| 8 | VideoChat-Flash | 2B | 42 | 0.60 | 0.63 | 0.58 |
| 9 | MiniCPM-V | 8B | 42 | 0.60 | 0.66 | 0.55 |
| 10 | Video-XL | 7B | 36 | 0.59 | 0.62 | 0.57 |
| 11 | ShareGPT4Video | 8B | 29 | 0.58 | 0.61 | 0.56 |
| 12 | InternVL2 | 8B | 27 | 0.58 | 0.62 | 0.55 |
| 13 | TimeChat | 7B | 20 | 0.56 | 0.59 | 0.54 |
| 14 | LLaVA-NeXT-Video | 7B | 16 | 0.55 | 0.56 | 0.54 |
| 15 | TS-LLaVA | 7B | 9 | 0.53 | 0.54 | 0.53 |
| 16 | Oryx | 7B | 6 | 0.52 | 0.54 | 0.52 |
| 17 | LongVA | 7B | 0 | 0.48 | 0.48 | 0.48 |
Per-judge breakdown
| Overall Rank | Model | GTE-Qwen2-7B β | KaLM-Gemma3-12B β | Nemo-8B β | NV-Embed-v2 β | Qwen3-8B β | Mean HM-CFCoarseFine-F1 | Borda |
|---|---|---|---|---|---|---|---|---|
| 1 | VideoLLaMA3 | 0.690.750.64 | 0.790.820.76 | 0.620.680.57 | 0.550.680.47 | 0.700.720.69 | 0.670.730.63 | 80 |
| 2 | mPLUG-Owl3 | 0.680.740.64 | 0.770.790.75 | 0.610.670.55 | 0.530.630.46 | 0.690.700.68 | 0.660.710.62 | 75 |
| 3 | LLaVA-OneVision | 0.650.700.62 | 0.760.790.74 | 0.590.650.54 | 0.520.640.44 | 0.670.680.67 | 0.640.690.60 | 67 |
| 4 | ViLAMP | 0.660.710.62 | 0.760.780.75 | 0.590.650.55 | 0.500.600.44 | 0.680.690.66 | 0.640.690.60 | 67 |
| 5 | LongVU | 0.640.700.60 | 0.750.780.73 | 0.560.630.51 | 0.510.610.44 | 0.670.670.67 | 0.630.680.59 | 61 |
| 6 | Qwen2.5-72B | 0.630.660.59 | 0.750.770.72 | 0.560.630.50 | 0.500.590.43 | 0.660.680.65 | 0.620.670.58 | 55 |
| 7 | Qwen2.5-32B | 0.610.660.57 | 0.740.770.72 | 0.540.600.49 | 0.490.580.42 | 0.650.660.65 | 0.610.650.57 | 48 |
| 8 | VideoChat-Flash | 0.610.630.59 | 0.730.740.72 | 0.550.590.52 | 0.450.540.40 | 0.660.670.64 | 0.600.630.58 | 42 |
| 9 | MiniCPM-V | 0.600.670.55 | 0.740.780.70 | 0.530.610.47 | 0.480.600.40 | 0.650.660.64 | 0.600.660.55 | 42 |
| 10 | Video-XL | 0.600.630.57 | 0.730.740.72 | 0.520.560.49 | 0.470.550.40 | 0.630.620.65 | 0.590.620.57 | 36 |
| 11 | ShareGPT4Video | 0.600.640.57 | 0.710.720.70 | 0.510.550.48 | 0.460.550.40 | 0.620.610.64 | 0.580.610.56 | 29 |
| 12 | InternVL2 | 0.600.630.56 | 0.720.740.71 | 0.520.560.49 | 0.450.570.38 | 0.620.610.63 | 0.580.620.55 | 27 |
| 13 | TimeChat | 0.570.580.55 | 0.710.720.71 | 0.500.520.49 | 0.420.530.35 | 0.600.590.62 | 0.560.590.54 | 20 |
| 14 | LLaVA-NeXT-Video | 0.540.550.53 | 0.700.690.71 | 0.480.490.48 | 0.420.500.37 | 0.590.560.63 | 0.550.560.54 | 16 |
| 15 | TS-LLaVA | 0.530.540.51 | 0.680.680.69 | 0.460.460.46 | 0.400.470.35 | 0.580.560.62 | 0.530.540.53 | 9 |
| 16 | Oryx | 0.510.520.51 | 0.670.670.68 | 0.460.470.45 | 0.400.480.34 | 0.580.560.61 | 0.520.540.52 | 6 |
| 17 | LongVA | 0.480.500.46 | 0.640.620.66 | 0.400.410.41 | 0.350.400.31 | 0.530.490.57 | 0.480.480.48 | 0 |
Among the 17 models evaluated in the paper, every model scores higher on coarse-grained than fine-grained similarity — the top-ranked model reaches 0.73 coarse against 0.63 fine, and the best mean HM-CF is 0.67: long-form video description is a long way from solved.
BibTeX
@misc{ali2026clipccbench,
title = {{CLIP-CC-Bench}: Evaluating Paragraph-Level Video Descriptions in Video--Language Models},
author = {Ali, Mukhtiar and Dubey, Harsh and Mishra, Sugam and Pack, Chulwoo},
year = {2026},
eprint = {2608.04302},
archivePrefix = {arXiv},
primaryClass = {cs.CV},
note = {Presented at the 2nd Workshop on Evaluation for Multimodal Generation (EvalMG), ACM SIGIR 2026},
url = {https://arxiv.org/abs/2608.04302}
}