LLM Leaderboard

A traceable LLM ranking snapshot that brings together multiple public benchmark leaderboards into one decision table, refreshed every day.

Models

20

Signals

44

Updated

2026-09-08

Refreshed daily at 00:00 HKT

Consensus snapshot

Major benchmark signals, normalized into one decision table.

#1Claude Fable 5.1Anthropic · Reasoning index #1 · Text preference #3 · Agent tasks #198
#2Claude Fable 5Anthropic · Reasoning index #4 · Text preference #1 · Agent tasks #393
#3GPT-6 AstraOpenAI · Reasoning index #290
#4Claude Opus 5Anthropic · Reasoning index #3 · Text preference #7 · Agent tasks #288

Consensus Ranking

The consensus score is not a raw benchmark score. It normalizes cross-benchmark positions and groups effort levels and variants of one product line under its best result.

Weight mixReasoning index 40%Text preference 30%Agent tasks 30%
Rank01Proprietary

Claude Fable 5.1

Anthropic

  • Reasoning index #1
  • Text preference #3
  • Agent tasks #1

98

100%

3 benchmark signals

Rank02Proprietary

Claude Fable 5

Anthropic

  • Reasoning index #4
  • Text preference #1
  • Agent tasks #3

93

100%

3 benchmark signals

Rank03Proprietary

GPT-6 Astra

OpenAI

  • Reasoning index #2

90

40%

1 benchmark signals

Rank04Proprietary

Claude Opus 5

Anthropic

  • Reasoning index #3
  • Text preference #7
  • Agent tasks #2

88

100%

3 benchmark signals

Rank05Proprietary

Muse Spark 1.3

Meta

  • Reasoning index #5

79

40%

1 benchmark signals

Rank06Proprietary

GPT-5.6 Sol

OpenAI

  • Reasoning index #6
  • Text preference #14
  • Agent tasks #4

73

100%

3 benchmark signals

Rank07Proprietary

Claude Opus 4.7

Anthropic

  • Text preference #4
  • Agent tasks #10

73

60%

2 benchmark signals

Rank08Open weights

Kimi K3

Kimi

  • Reasoning index #10
  • Text preference #10
  • Agent tasks #6

69

100%

3 benchmark signals

Rank09Proprietary

Claude Opus 4.6

Anthropic

  • Text preference #2
  • Agent tasks #17

63

60%

2 benchmark signals

Rank10Open weights

Hy4 preview

Tencent

  • Agent tasks #9

63

30%

1 benchmark signals

Rank11Proprietary

Claude Opus 4.8

Anthropic

  • Text preference #15
  • Agent tasks #5

61

60%

2 benchmark signals

Rank12Proprietary

Gemini 3.8 Flash

Google

  • Reasoning index #14
  • Text preference #6
  • Agent tasks #14

58

100%

3 benchmark signals

Rank13Proprietary

Muse Spark

Meta

  • Text preference #11

56

30%

1 benchmark signals

Rank14Proprietary

Qwen3.8 Max

Alibaba

  • Reasoning index #7
  • Text preference #19
  • Agent tasks #13

54

100%

3 benchmark signals

Rank15Proprietary

GPT 5.5

OpenAI

  • Text preference #16
  • Agent tasks #8

54

60%

2 benchmark signals

Rank16Open weights

GLM-5.3

Z AI

  • Reasoning index #8
  • Text preference #17
  • Agent tasks #18

49

100%

3 benchmark signals

Rank17Proprietary

Gemini 3 Pro

Google

  • Text preference #13

48

30%

1 benchmark signals

Rank18Proprietary

Muse Spark 1.2

Meta

  • Reasoning index #16
  • Text preference #5
  • Agent tasks #28

41

100%

3 benchmark signals

Rank19Open weights

Qwen3.8 2.4T A95B

Alibaba

  • Reasoning index #15

41

40%

1 benchmark signals

Rank20Proprietary

Grok 4.6

SpaceXAI

  • Reasoning index #9
  • Text preference #42
  • Agent tasks #16

39

100%

3 benchmark signals

Methodology

Weighted consensus, not benchmark shopping.

  1. 01Three public benchmark leaderboards are fetched automatically every day at 00:00 HKT and cached for 24 hours.
  2. 02Each benchmark rank is converted into position points, then averaged with the declared weight mix.
  3. 03Effort levels and variants of one product line are grouped, keeping only its best result per benchmark.
  4. 04Models missing some benchmark coverage receive a small coverage discount, so a single leaderboard spike does not dominate the table.
  5. 05The ranking is a procurement and experimentation entry point, not an absolute answer. Cost, latency, context, data policy, and task-level testing still matter.

Turn rankings into a practical model selection process.

If you want to connect this leaderboard to model selection, prompt tests, or agent workflows, we can turn it into a repeatable evaluation lane.

Discuss AI workflow