🇩🇪 MTEB-DE — German Embedding Benchmark

A reproducible leaderboard across four task types — retrieval, reranking, semantic textual similarity (STS), and clustering — for German / multilingual embedding models, with a net-new hard jobs retrieval set derived from official-API job data.

Click a column header (or use the controls below) to sort. Overall is the mean of a model's available task scores, so models evaluated on fewer tasks are not directly comparable on Overall — sort by a task column for a fair head-to-head.

Sort by
Order

Tasks & metrics

Task Config Metric Source
Retrieval germandpr nDCG@10 mteb/GermanDPR (CC-BY-4.0)
Retrieval jobs nDCG@10 derived from mischeiwiller/german-job-postings (CC-BY-4.0)
Reranking jobs nDCG@10 mined hard negatives over the jobs set
STS stsb_de Spearman PhilipMay/stsb_multi_mt (de), loaded upstream
Clustering jobs_occupation V-measure derived job title → ESCO occupation

All scores come from one documented harness command per task; every model and data source is pinned to a Hub revision for byte-reproducible reruns.

Caveats (v1): the reranking set is small and was mined with a single encoder (miner bias — see the suite card / OPEN-QUESTIONS); a fine-tuned bge-reranker-de and an enlarged, model-agnostic reranking set are in progress.

📦 Suite: mischeiwiller/mteb-de