Benchmarks
Model performance scores across popular evaluation benchmarks.
GPT-5
90.2
Quelle: -
Claude Opus 4
90.1
Quelle: -
DeepSeek R1
89.1
Quelle: -
Gemini 2.5 Pro
89
Quelle: -
GPT-4o
83.1
Quelle: -
Model performance scores across popular evaluation benchmarks.