🇩🇪 MTEB-DE — German Embedding Benchmark
A reproducible leaderboard across four task types — retrieval, reranking, semantic textual similarity (STS), and clustering — for German / multilingual embedding models, with a net-new hard jobs retrieval set derived from official-API job data.
Click a column header (or use the controls below) to sort. Overall is the mean of a model's available task scores, so models evaluated on fewer tasks are not directly comparable on Overall — sort by a task column for a fair head-to-head.
Tasks & metrics
| Task | Config | Metric | Source |
|---|---|---|---|
| Retrieval | germandpr |
nDCG@10 | mteb/GermanDPR (CC-BY-4.0) |
| Retrieval | jobs |
nDCG@10 | derived from mischeiwiller/german-job-postings (CC-BY-4.0) |
| Reranking | jobs |
nDCG@10 | mined hard negatives over the jobs set |
| STS | stsb_de |
Spearman | PhilipMay/stsb_multi_mt (de), loaded upstream |
| Clustering | jobs_occupation |
V-measure | derived job title → ESCO occupation |
All scores come from one documented harness command per task; every model and data source is pinned to a Hub revision for byte-reproducible reruns.
Caveats (v1): the reranking set is small and was mined with a single encoder
(miner bias — see the suite card / OPEN-QUESTIONS); a fine-tuned bge-reranker-de
and an enlarged, model-agnostic reranking set are in progress.
📦 Suite: mischeiwiller/mteb-de