AGIラボラボメン

ベンチマーク

API料金

14 件のベンチマーク

知識11 モデル · 11 公表値

GPQA Diamond

Diamond split

公表値の高い順・評価条件は異なります · %

評価条件と原典を確認する

専門家が作成した科学の難問で、知識と推論を測る。

Diamond分割を使用。ツール、推論設定、集計方法を区別します。

GPQA

評価条件 1

Diamond / No tools / xhigh / Accuracy · trial count not specified / OpenAI research environment · API

GPT-5.2 Thinking · 92.4%

OpenAI · openai:gpt-5.2

開発元の公表値

資料公表 2025-12-11 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Appendix → Academic → GPQA Diamond; evaluation footnote

一部の評価条件は未公表です。

評価条件 2

Diamond / No tools / xhigh / Accuracy · trial count not specified / OpenAI research environment · API

GPT-5.2 Pro · 93.2%

OpenAI · openai:gpt-5.2-pro

開発元の公表値

資料公表 2025-12-11 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Appendix → Academic → GPQA Diamond; evaluation footnote

一部の評価条件は未公表です。

評価条件 3

Diamond / No tools / high / Accuracy · trial count not specified / OpenAI research environment · API

GPT-5.1 Thinking · 88.1%

OpenAI · openai:gpt-5.1

開発元の公表値

資料公表 2025-12-11 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Appendix → Academic → GPQA Diamond; evaluation footnote

一部の評価条件は未公表です。

評価条件 4

Diamond / Not specified in source / Chain of thought / 0-shot · accuracy / Meta instruction-tuned evaluation · English

Llama 3.3 70B Instruct · 50.5%

Meta · meta-llama/Llama-3.3-70B-Instruct

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Benchmarks — English Text → GPQA Diamond (CoT)

一部の評価条件は未公表です。

評価条件 5

Diamond / Not specified in source / Chain of thought / 0-shot · accuracy / Meta instruction-tuned evaluation · English

Llama 3.1 405B Instruct · 49%

Meta · meta-llama/Meta-Llama-3.1-405B-Instruct

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Benchmarks — English Text → GPQA Diamond (CoT)

一部の評価条件は未公表です。

評価条件 6

Diamond / Not specified in source / Chain of thought / 5-shot · trial count not specified / Mistral provider evaluation · text

Mistral Small 3.2 24B Instruct · 46.13%

Mistral AI · mistralai/Mistral-Small-3.2-24B-Instruct-2506

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Benchmark Results → Text → STEM → GPQA Diamond

一部の評価条件は未公表です。

評価条件 7

Diamond / Not specified in source / Thinking / pass@1 · trial count not specified / DeepSeek research evaluation · temperature 1.0 · 128k context

DeepSeek V3.2 Thinking · 82.4%

DeepSeek · deepseek-ai/DeepSeek-V3.2

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

§4.1 Table 2 → GPQA Diamond

一部の評価条件は未公表です。

評価条件 8

GPQA Diamond; dataset revision notreported / notreported / Maximum score across effort settings; exact effort notreported / notreported / OpenAI research environment or API; exact harness notreported

GPT-6 Astra · 96%

OpenAI · openai:gpt-6-astra

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Academic table: GPQA Diamond row, GPT-6 Astra column; evaluation configuration beneath Abstract reasoning table

一部の評価条件は未公表です。

評価条件 9

GPQA Diamond; dataset revision notreported / No tools / High thinking / Single attempt (pass@1); default sampling; no majority voting or parallel test-time compute; exact trial count notreported / Gemini API; exact harness version notreported

Gemini 3.1 Flash Lite Preview · 86.9%

Google · google:gemini-3.1-flash-lite-preview

開発元の公表値

資料公表 2026-03-03 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation > Results (as of March 2026) > GPQA Diamond Scientific knowledge > Gemini 3.1 Flash-Lite High column; No tools row note

一部の評価条件は未公表です。

評価手法PDFがmodel-id gemini-3.1-flash-lite-previewを明記。3月のPreview評価であり、現行stable版への転記は行わない。

評価条件 10

Diamond · 198 questions / Tool setting not specified in §2.9 / Adaptive thinking · max effort / Mean over 10 trials / Anthropic evaluation · default temperature/top_p

Claude Sonnet 4.6 · 89.9%

Anthropic · anthropic:claude-sonnet-4-6

開発元の公表値

資料公表 2026-02-17 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

System Card §2.9, pp.21–22

一部の評価条件は未公表です。

評価条件 11

Diamond / No tools / Thinking (High) / pass@1 · mean over trials / Gemini API · default sampling

Gemini 3.1 Pro Preview · 94.3%

Google · google:gemini-3.1-pro-preview

開発元の公表値

資料公表 2026-02-19 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation → GPQA Diamond → Gemini 3.1 Pro Thinking (High)

数学7 モデル · 8 公表値

AIME 2025

2025

公表値の高い順・評価条件は異なります · %

評価条件と原典を確認する

2025年の数学競技問題。年の違う結果と分けて読む。

同じ年でもI・II、pass@1、多数決では評価条件が異なります。

AIME 2025 dataset

評価条件 1

AIME 2025 · I/II coverage not specified / No tools / xhigh / Accuracy · trial count not specified / OpenAI research environment · API

GPT-5.2 Thinking · 100%

OpenAI · openai:gpt-5.2

開発元の公表値

資料公表 2025-12-11 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Appendix → Academic → AIME 2025; evaluation footnote

一部の評価条件は未公表です。

評価条件 2

AIME 2025 · I/II coverage not specified / No tools / xhigh / Accuracy · trial count not specified / OpenAI research environment · API

GPT-5.2 Pro · 100%

OpenAI · openai:gpt-5.2-pro

開発元の公表値

資料公表 2025-12-11 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Appendix → Academic → AIME 2025; evaluation footnote

一部の評価条件は未公表です。

評価条件 3

AIME 2025 · I/II coverage not specified / No tools / high / Accuracy · trial count not specified / OpenAI research environment · API

GPT-5.1 Thinking · 94%

OpenAI · openai:gpt-5.1

開発元の公表値

資料公表 2025-12-11 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Appendix → Academic → AIME 2025; evaluation footnote

一部の評価条件は未公表です。

評価条件 4

AIME 2025 · I/II coverage not specified / Not specified in source / Thinking · output limit 81,920 / Trial count not specified / Qwen provider evaluation · sampling details not fully specified

Qwen3 235B A22B Thinking 2507 · 92.3%

Alibaba · Qwen/Qwen3-235B-A22B-Thinking-2507

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Performance → Reasoning → AIME25; footnote &

一部の評価条件は未公表です。

評価条件 5

AIME 2025 · I/II coverage not specified / No tools / Thinking · 96k budget / Mean of 32 runs / Moonshot evaluation · INT4 · temperature 1.0 · 256k context

Kimi K2 Thinking · 94.5%

Moonshot AI · moonshotai/Kimi-K2-Thinking

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation Results → AIME25 no tools; note 2.3

一部の評価条件は未公表です。

評価条件 6

AIME 2025 · I/II coverage not specified / Python / Thinking · 96k budget / Mean of 16 runs / Moonshot evaluation · INT4 · temperature 1.0 · 256k context

Kimi K2 Thinking · 99.1%

Moonshot AI · moonshotai/Kimi-K2-Thinking

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation Results → AIME25 w/ python; note 2.3

一部の評価条件は未公表です。

評価条件 7

AIME 2025 · I/II coverage not specified / Not specified in source / Thinking · step-by-step boxed-answer prompt / pass@1 · trial count not specified / DeepSeek research evaluation · temperature 1.0 · 128k context

DeepSeek V3.2 Thinking · 93.1%

DeepSeek · deepseek-ai/DeepSeek-V3.2

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

§4.1 Table 2 → AIME 2025; evaluation prompt

一部の評価条件は未公表です。

評価条件 8

AIME 2025 · I/II coverage not specified / No tools / Adaptive thinking · max effort / Mean over 10 trials / Anthropic evaluation · default temperature/top_p

Claude Sonnet 4.6 · 95.6%

Anthropic · anthropic:claude-sonnet-4-6

開発元の公表値

資料公表 2026-02-17 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

System Card §2.10, p.22

一部の評価条件は未公表です。

Anthropicは、学習データへの問題混入でスコアが高くなった可能性を指摘しています。

知識6 モデル · 6 公表値

MMLU-Pro

Pro

公表値の高い順・評価条件は異なります · %

評価条件と原典を確認する

知識に加え、より難しい推論を求めるMMLUの拡張。

元のMMLUとは別の問題集合。CoTやfew-shotの条件を確認します。

TIGER-AI-Lab

評価条件 1

MMLU-Pro / Not specified in source / Chain of thought / 5-shot · macro-average accuracy / Meta instruction-tuned evaluation · English

Llama 3.3 70B Instruct · 68.9%

Meta · meta-llama/Llama-3.3-70B-Instruct

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Benchmarks — English Text → MMLU Pro (CoT)

一部の評価条件は未公表です。

評価条件 2

MMLU-Pro / Not specified in source / Chain of thought / 5-shot · macro-average accuracy / Meta instruction-tuned evaluation · English

Llama 3.1 405B Instruct · 73.3%

Meta · meta-llama/Meta-Llama-3.1-405B-Instruct

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Benchmarks — English Text → MMLU Pro (CoT)

一部の評価条件は未公表です。

評価条件 3

MMLU-Pro / Not specified in source / Chain of thought / 5-shot · trial count not specified / Mistral provider evaluation · text

Mistral Small 3.2 24B Instruct · 69.06%

Mistral AI · mistralai/Mistral-Small-3.2-24B-Instruct-2506

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Benchmark Results → Text → STEM → MMLU Pro

一部の評価条件は未公表です。

評価条件 4

MMLU-Pro / Not specified in source / Thinking · output limit 32,768 / Trial count not specified / Qwen provider evaluation · sampling details not fully specified

Qwen3 235B A22B Thinking 2507 · 84.4%

Alibaba · Qwen/Qwen3-235B-A22B-Thinking-2507

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Performance → Knowledge → MMLU-Pro; footnote &

一部の評価条件は未公表です。

評価条件 5

MMLU-Pro / No tools / Thinking · token budget not specified / Trial count not specified / Moonshot evaluation · INT4 · temperature 1.0 · 256k context

Kimi K2 Thinking · 84.6%

Moonshot AI · moonshotai/Kimi-K2-Thinking

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation Results → General → MMLU-Pro; note 2.1

一部の評価条件は未公表です。

評価条件 6

MMLU-Pro / Not specified in source / Thinking / Exact match · trial count not specified / DeepSeek research evaluation · temperature 1.0 · 128k context

DeepSeek V3.2 Thinking · 85%

DeepSeek · deepseek-ai/DeepSeek-V3.2

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

§4.1 Table 2 → MMLU-Pro

一部の評価条件は未公表です。

知識4 モデル · 6 公表値

Humanity’s Last Exam

Full set · text + multimodal

公表値の高い順・評価条件は異なります · %

評価条件と原典を確認する

幅広い専門分野の最先端の学術問題に答える。

全文・画像を含む評価。テキストのみやHLE-Rolling、ツールの有無を混ぜません。

Center for AI Safety / Scale AI

評価条件 1

Full multimodal set; 2,500 questions / None / Adaptive thinking, max effort; auto; total token cap 1M; no compaction / 5-trial average; default temperature/top_p / Anthropic evaluation; Claude Opus 4.6 grader; harness version notreported

Claude Fable 5.1 · 60.9%

Anthropic · anthropic:claude-fable-5-1

開発元の公表値

資料公表 2026-09-01 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Table 8.1.A p.167, configuration caption pp.167-168; section 8.12.1 p.176; corresponding Fable 5.1 column in official announcement

一部の評価条件は未公表です。

本番safeguardsを有効にした公表値。公式発表では介入時に旧Opusモデルへfallbackする評価構成が記されている。純粋な単一モデル実行との比較は未確認。

評価条件 2

Full multimodal set; 2,500 questions / Web search, restricted web fetch, programmatic tool calling, code execution; HLE source blocklist; contamination review / Adaptive thinking, max effort; auto; total token cap 1M; no compaction / 5-trial average; default temperature/top_p / Anthropic evaluation; Claude Opus 4.6 grader; harness version notreported

Claude Fable 5.1 · 65%

Anthropic · anthropic:claude-fable-5-1

開発元の公表値

資料公表 2026-09-01 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Table 8.1.A p.167, configuration caption pp.167-168; section 8.12.1 p.176; corresponding Fable 5.1 column in official announcement

一部の評価条件は未公表です。

本番safeguardsを有効にした公表値。公式発表では介入時に旧Opusモデルへfallbackする評価構成が記されている。純粋な単一モデル実行との比較は未確認。

評価条件 3

Full set; text + multimodal; dataset revision notreported / No tools / High thinking / Single attempt (pass@1); default sampling; no majority voting or parallel test-time compute; exact trial count notreported / Gemini API; exact harness version notreported

Gemini 3.1 Flash Lite Preview · 16%

Google · google:gemini-3.1-flash-lite-preview

開発元の公表値

資料公表 2026-03-03 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation > Results (as of March 2026) > Humanity’s Last Exam Academic reasoning (full set, text + MM) > Gemini 3.1 Flash-Lite High column; No tools row note

一部の評価条件は未公表です。

評価手法PDFがmodel-id gemini-3.1-flash-lite-previewを明記。3月のPreview評価であり、現行stable版への転記は行わない。

評価条件 4

Full set · 2,500 questions · multimodal / No tools / Adaptive thinking · max effort / Mean over 10 trials / Anthropic evaluation · Claude Sonnet 4.5 grader

Claude Sonnet 4.6 · 33.2%

Anthropic · anthropic:claude-sonnet-4-6

開発元の公表値

資料公表 2026-02-17 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

System Card Table 2.1.A and §2.20.2

評価条件 5

Full set · text + multimodal / No tools / Thinking (High) / pass@1 · mean over trials / Gemini API · default sampling

Gemini 3.1 Pro Preview · 44.4%

Google · google:gemini-3.1-pro-preview

開発元の公表値

資料公表 2026-02-19 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation → Humanity’s Last Exam → No tools

評価条件 6

Full set · text + multimodal / Search (benchmark blocklist) + Code / Thinking (High) / pass@1 · mean over trials / Gemini API · default sampling

Gemini 3.1 Pro Preview · 51.4%

Google · google:gemini-3.1-pro-preview

開発元の公表値

資料公表 2026-02-19 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation → Humanity’s Last Exam → Search (blocklist) + Code

知識3 モデル · 5 公表値

Humanity’s Last Exam

Text-only subset

公表値の高い順・評価条件は異なります · %

評価条件と原典を確認する

HLEのうち、画像を必要としない問題だけを評価する。

全文・画像を含む結果とは別に掲載。ツール利用と問題集合の版を結果ごとに確認します。

Center for AI Safety / Scale AI

評価条件 1

Text-only HLE / Not specified in source / Thinking · output limit 81,920 / Trial count not specified / Qwen provider evaluation · sampling details not fully specified

Qwen3 235B A22B Thinking 2507 · 18.2%

Alibaba · Qwen/Qwen3-235B-A22B-Thinking-2507

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Performance → HLE; footnotes # and &

一部の評価条件は未公表です。

評価条件 2

Text-only HLE / No tools / Thinking · 96k budget / Trial count not specified / Moonshot evaluation · INT4 · temperature 1.0 · 256k context

Kimi K2 Thinking · 23.9%

Moonshot AI · moonshotai/Kimi-K2-Thinking

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation Results → Reasoning → HLE (Text-only); notes 2.1–2.2

一部の評価条件は未公表です。

評価条件 3

Text-only HLE / Search + code interpreter + browsing · Hugging Face blocked / 120 steps · 48k reasoning per step / o3-mini judge · trial count not specified / Moonshot evaluation · INT4 · temperature 1.0 · 256k context

Kimi K2 Thinking · 44.9%

Moonshot AI · moonshotai/Kimi-K2-Thinking

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation Results → HLE w/ tools; notes 4.1–4.6

一部の評価条件は未公表です。

評価条件 4

Text-only HLE / Search + code interpreter + browsing · Hugging Face blocked / Heavy mode / 8 parallel trajectories then aggregation / Moonshot evaluation · INT4 · temperature 1.0 · 256k context

Kimi K2 Thinking · 51%

Moonshot AI · moonshotai/Kimi-K2-Thinking

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation Results → HLE heavy; notes 4 and 6

一部の評価条件は未公表です。

評価条件 5

Text-only HLE / Not specified in source / Thinking · step-by-step boxed-answer prompt / pass@1 · trial count not specified / DeepSeek research evaluation · temperature 1.0 · 128k context

DeepSeek V3.2 Thinking · 25.1%

DeepSeek · deepseek-ai/DeepSeek-V3.2

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

§4.1 Table 2 → HLE; evaluation prompt

一部の評価条件は未公表です。

同じ資料で、HLE公式プロンプトを使った別条件の値は23.9と報告されています。

推論3 モデル · 3 公表値

ARC-AGI-2

2 · 2025

公表値の高い順・評価条件は異なります · %

評価条件と原典を確認する

少数の例から、図形の規則を見つけて応用する。

公開・準非公開・非公開の評価集合、試行数、計算費用を区別します。

ARC Prize

評価条件 1

ARC-AGI-2 Verified · subset not specified / Not specified in source / xhigh / Verified accuracy · attempts not specified / OpenAI research environment · API

GPT-5.2 Thinking · 52.9%

OpenAI · openai:gpt-5.2

開発元の公表値

資料公表 2025-12-11 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Appendix → Abstract reasoning → ARC-AGI-2 (Verified); evaluation footnote

一部の評価条件は未公表です。

評価条件 2

ARC-AGI-2 semi-private dataset / notreported / Max effort / notreported / ARC Prize Foundation; version notreported

Claude Fable 5.1 · 90%

Anthropic · anthropic:claude-fable-5-1

開発元の公表値

資料公表 2026-09-01 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Section 8.16 p.196; Table 8.1.A p.167

一部の評価条件は未公表です。

Anthropicのカードが報告するARC Prizeの検証値。ARC側の原典は今回未取得。

評価条件 3

ARC-AGI-2; split notreported / notreported / Maximum score across effort settings; exact effort notreported / notreported / OpenAI research environment or API; exact harness notreported

GPT-6 Astra · 95%

OpenAI · openai:gpt-6-astra

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Abstract reasoning table: ARC-AGI-2 row, GPT-6 Astra column

一部の評価条件は未公表です。

推論3 モデル · 3 公表値

ARC-AGI-1

1 · 2019

公表値の高い順・評価条件は異なります · %

評価条件と原典を確認する

図形の入出力から規則を推測する、最初のARC評価。

評価集合と推論時の計算量を明記。新しい版とは分けて扱います。

ARC Prize

評価条件 1

ARC-AGI-1 Verified · subset not specified / Not specified in source / xhigh / Verified accuracy · attempts not specified / OpenAI research environment · API

GPT-5.2 Thinking · 86.2%

OpenAI · openai:gpt-5.2

開発元の公表値

資料公表 2025-12-11 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Appendix → Abstract reasoning → ARC-AGI-1 (Verified); evaluation footnote

一部の評価条件は未公表です。

評価条件 2

ARC-AGI-1 semi-private dataset / notreported / Max effort / notreported / ARC Prize Foundation; version notreported

Claude Fable 5.1 · 97.5%

Anthropic · anthropic:claude-fable-5-1

開発元の公表値

資料公表 2026-09-01 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Section 8.16 p.196; Table 8.1.A p.167

一部の評価条件は未公表です。

Anthropicのカードが報告するARC Prizeの検証値。ARC側の原典は今回未取得。

評価条件 3

ARC-AGI-1; split notreported / notreported / Maximum score across effort settings; exact effort notreported / notreported / OpenAI research environment or API; exact harness notreported

GPT-6 Astra · 98.5%

OpenAI · openai:gpt-6-astra

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Abstract reasoning table: ARC-AGI-1 row, GPT-6 Astra column

一部の評価条件は未公表です。

知識3 モデル · 3 公表値

MMLU

Original

公表値の高い順・評価条件は異なります · %

評価条件と原典を確認する

幅広い分野の知識を問う、多肢選択式の評価。

プロンプトの例示数と集計方法を確認。MMLU-Proや多言語版とは分けます。

MMLU authors

評価条件 1

MMLU / Not specified in source / Chain of thought / 0-shot · macro-average accuracy / Meta instruction-tuned evaluation · English

Llama 3.3 70B Instruct · 86%

Meta · meta-llama/Llama-3.3-70B-Instruct

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Benchmarks — English Text → MMLU (CoT)

一部の評価条件は未公表です。

評価条件 2

MMLU / Not specified in source / Chain of thought / 0-shot · macro-average accuracy / Meta instruction-tuned evaluation · English

Llama 3.1 405B Instruct · 88.6%

Meta · meta-llama/Meta-Llama-3.1-405B-Instruct

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Benchmarks — English Text → MMLU (CoT)

一部の評価条件は未公表です。

評価条件 3

MMLU / Not specified in source / Not specified in source / Shot count and aggregation not specified / Mistral provider evaluation · text

Mistral Small 3.2 24B Instruct · 80.5%

Mistral AI · mistralai/Mistral-Small-3.2-24B-Instruct-2506

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Benchmark Results → Text → STEM → MMLU

一部の評価条件は未公表です。

コード2 モデル · 2 公表値

LiveCodeBench

release_v6 · code generation

公表値の高い順・評価条件は異なります · %

評価条件と原典を確認する

競技プログラミングの問題でコード生成を測る。

v6は2023年5月〜2025年4月の問題。日付フィルタとpass@kを区別します。

LiveCodeBench

評価条件 1

v6 · 2025-02–2025-05 as reported by Qwen / Not specified in source / Thinking · output limit 81,920 / Trial count not specified / Qwen provider evaluation · sampling details not fully specified

Qwen3 235B A22B Thinking 2507 · 74.1%

Alibaba · Qwen/Qwen3-235B-A22B-Thinking-2507

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Performance → Coding → LiveCodeBench v6 (25.02-25.05)

一部の評価条件は未公表です。

公表資料の日付範囲をそのまま記載しています。v6全体や他の日付範囲とは別の評価です。

評価条件 2

LiveCodeBench v6 · date filter not specified / No tools / Thinking · 128k budget / Mean of 5 runs / Moonshot evaluation · INT4 · temperature 1.0 · 256k context

Kimi K2 Thinking · 83.1%

Moonshot AI · moonshotai/Kimi-K2-Thinking

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Evaluation Results → Coding → LiveCodeBenchV6; notes 2.2 and 5.3

一部の評価条件は未公表です。

推論1 モデル · 1 公表値

ARC-AGI-3

3 · interactive

公表値の高い順・評価条件は異なります · %

評価条件と原典を確認する

未知の環境を探索し、行動しながら学ぶ能力を測る。

人の行動効率を基準にした評価。ARC-AGI-1・2の正答率とは別の指標です。

ARC Prize

評価条件 1

ARC-AGI-3 interactive; dataset revision notreported / notreported / Maximum score across effort settings; exact effort notreported / notreported / OpenAI research environment or API; exact harness notreported

GPT-6 Astra · 99.9%

OpenAI · openai:gpt-6-astra

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Abstract reasoning table: ARC-AGI-3 row, GPT-6 Astra column

一部の評価条件は未公表です。

総合1 モデル · 1 公表値

LiveBench

2024-11-25 release

公表値の高い順・評価条件は異なります · score

評価条件と原典を確認する

2024年11月版の問題で測った、当時のモデルの総合性能。

2025年4月版とは分離。カテゴリ範囲や推論設定を公表資料に沿って記載します。

LiveBench

評価条件 1

LiveBench 20241125 / Not specified in source / Thinking · output limit 81,920 / Category aggregation not specified / Qwen provider evaluation · sampling details not fully specified

Qwen3 235B A22B Thinking 2507 · 78.4 score

Alibaba · Qwen/Qwen3-235B-A22B-Thinking-2507

開発元の公表値

資料公表 日付未確認 · 収集・原典確認 2026-09-08 · 測定日 未公表

スコアの原典評価方法

Performance → Reasoning → LiveBench 20241125; footnote &

一部の評価条件は未公表です。

スコア収集前の評価

LiveBench2025-04-25 release評価の公式資料
FrontierMathTiers 1–3 · v2評価の公式資料

スコア未収録は、0点や未測定を意味しません。

ベンチマークの版・評価条件ごとに掲載しています。スコア未収録は、未測定や0点を意味しません。網羅的なランキングではなく、原典で確認できた公表値を段階的に収録しています。

API料金

公式価格あり 7モデル公開時価格 1モデル

公式API価格。USD、100万トークンあたり。標準枠・短い入力・キャッシュなし。日付は各価格の確認日。提供終了したPreviewには公開時価格の注記があります。
モデル入力 / 1M tokens出力 / 1M tokens

Claude Fable 5.1

Anthropic

$10.00公式出典

2026-09-08

掲載箇所

Model pricing: Claude Fable 5.1 row, Base input tokens / Output tokens; MTok definition; Long context pricing; Data residency pricing. First-party Claude API, standard text base input and text output, global routing, up to 1M context; excludes caching, Batch, fast mode, residency premiums and tool charges. Verified official snapshot observed 2026-09-08T00:33:51.031955+00:00.

$50.00公式出典

2026-09-08

掲載箇所

Model pricing: Claude Fable 5.1 row, Base input tokens / Output tokens; MTok definition; Long context pricing; Data residency pricing. First-party Claude API, standard text base input and text output, global routing, up to 1M context; excludes caching, Batch, fast mode, residency premiums and tool charges. Verified official snapshot observed 2026-09-08T00:33:51.031955+00:00.

GPT-6 Astra

OpenAI

$10.00公式出典

2026-09-08

掲載箇所

Pricing > Text tokens > Per 1M tokens > Input / Output; immediately following long-context and service-tier notes. OpenAI API Standard text tokens; uncached input; requests with input at or below 272K tokens. Above 272K input, full-request input rate is 2x and output rate is 1.5x. Excludes cache writes/reads, Batch, Flex, Fast and tool fees. Verified official snapshot observed 2026-09-08T00:33:51.031955+00:00.

$50.00公式出典

2026-09-08

掲載箇所

Pricing > Text tokens > Per 1M tokens > Input / Output; immediately following long-context and service-tier notes. OpenAI API Standard text tokens; uncached input; requests with input at or below 272K tokens. Above 272K input, full-request input rate is 2x and output rate is 1.5x. Excludes cache writes/reads, Batch, Flex, Fast and tool fees. Verified official snapshot observed 2026-09-08T00:33:51.031955+00:00.

Grok 4.3

SpaceXAI

$1.25公式出典

2026-09-22

掲載箇所

$.prompt_text_token_price

$2.50公式出典

2026-09-22

掲載箇所

$.completion_text_token_price

Grok 4.5

SpaceXAI

$2.00公式出典

2026-09-22

掲載箇所

$.prompt_text_token_price

$6.00公式出典

2026-09-22

掲載箇所

$.completion_text_token_price

Grok 4.6

SpaceXAI

$2.00公式出典

2026-09-22

掲載箇所

$.prompt_text_token_price

$6.00公式出典

2026-09-22

掲載箇所

$.completion_text_token_price

Grok 4.7

SpaceXAI

$2.00公式出典

2026-09-22

掲載箇所

$.prompt_text_token_price

$6.00公式出典

2026-09-22

掲載箇所

$.completion_text_token_price

xAI: Grok Build 0.1

SpaceXAI

$1.00公式出典

2026-09-22

掲載箇所

$.prompt_text_token_price

$2.00公式出典

2026-09-22

掲載箇所

$.completion_text_token_price

Gemini 3.1 Flash Lite Preview

Google

Preview公開時(2026-03-03)の価格。2026-05-25に提供終了。現在は利用できません。

$0.25公式出典

2026-09-08

掲載箇所

Evaluation > Results (March 2026) > Gemini 3.1 Flash-Lite High column > Input price $/1M tokens, no caching / Output price $/1M tokens; corroborated by Preview launch Cost-efficiency without compromise. Historical standard/base text API token rates at the March 3, 2026 Preview launch; input without caching; text output. Not a current purchasable Preview offer: official Preview docs report shutdown on May 25, 2026. Excludes Batch, Flex, Priority, caching and separate tool fees; stable-model pricing not substituted. Historical Preview launch: 2026-03-03; service ended: 2026-05-25. Source retrieved 2026-09-08T00:39:21.068483+00:00.

$1.50公式出典

2026-09-08

掲載箇所

Evaluation > Results (March 2026) > Gemini 3.1 Flash-Lite High column > Input price $/1M tokens, no caching / Output price $/1M tokens; corroborated by Preview launch Cost-efficiency without compromise. Historical standard/base text API token rates at the March 3, 2026 Preview launch; input without caching; text output. Not a current purchasable Preview offer: official Preview docs report shutdown on May 25, 2026. Excludes Batch, Flex, Priority, caching and separate tool fees; stable-model pricing not substituted. Historical Preview launch: 2026-03-03; service ended: 2026-05-25. Source retrieved 2026-09-08T00:39:21.068483+00:00.

USD / 100万トークン。標準枠・短い入力・キャッシュなしの確認済み価格です。提供終了したPreviewは公開時の価格を注記して掲載しています。短い入力の上限は各社で異なります。日付は原典の確認日。バッチ・長文・画像・音声などの料金は公式ページを確認してください。