🤖 AI benchmark: hit-rate of 7 models

Prematch En direct (in-play)

Seven external AI models (Hermes contour) independently analyze the same matchs — predicting the résultat (1X2), total (Plus de/Moins de), both équipes to score (BTTS) and the exact score. Here we honestly compare their pronostics against the real result after the final whistle and combine everything into a single accuracy rating. An informational and analytical snapshot, not betting advice.

⚠️ Data is still accumulating — counting starts from 09.07.2026, so all models are compared on the same events (early test pronostics are excluded). The sample is still small and not representative. Right now the snapshot holds 416 match(es), 1529 settled AI pronostics (Basketball). The figures below are N, not «a percentage you can trust»: the more matchs are played out, the more reliable the snapshot becomes. We show it transparently from day one, not only once the sample becomes «convenient».

Leaderboard · Basketball

Model N (settled) 1X2 Total points Exact score Composite accuracy
Claude
355 63.4%(225/355) 53.0%(141/266) 0.4%(1/270) 41.2%(367/891)
Kimi
195 63.6%(124/195) 51.5%(84/163) 0.0%(0/163) 39.9%(208/521)
Google AI
445 63.7%(283/444) 48.1%(199/414) 0.0%(0/420) 37.7%(482/1278)
ChatGPT
76 60.5%(46/76) 42.4%(25/59) 0.0%(0/59) 36.6%(71/194)
GLM 5.2
224 63.8%(143/224) 42.5%(94/221) 0.9%(2/222) 35.8%(239/667)
Qwen
197 64.0%(126/197) 35.8%(63/176) 0.6%(1/176) 34.6%(190/549)
DeepSeek
37 51.4%(19/37) 38.2%(13/34) 0.0%(0/37) 29.6%(32/108)

grey — sample <5, not representative; «—» — the modèle has not made a settled pronostic yet.

Composite accuracy — the share of correct pronostics sur all affiché marchés together: (sum of correct pronostics) ÷ (sum of all settled pronostics) sur the marchés 1X2 + Total points + Exact score. Each marché-pronostic weighs equally. This is hit-rate, not profitability — for money/ROI by modèle see /ai-agent. Total: a push (score exactly sur la cote) is excluded from the denominator. «Exact score» — the full final score was guessed correctly (H and A matched); pronostics with no recognized score do not count toward the denominator.

Composite modèle rating · all marchés · Basketball

Bar height = the modèle's composite accuracy sur all applicable marchés on the current sample. Sorted from best to worst.

41.2% (367/891)
Opus 4.8
39.9% (208/521)
Kimi 2.6
37.7% (482/1278)
Gemini 3.5 Flash
36.6% (71/194)
GPT 5.5
35.8% (239/667)
GLM 5.2
34.6% (190/549)
Qwen 3.7 Plus
29.6% (32/108)
DeepSeek V4 Pro

Bars are AI models by version; grey/dimmed — sample <5, not representative. The snapshot is informational, not betting advice.

Accuracy by marché · Basketball

Where each modèle is strong: one mini-bar per applicable marché, with the percentage and (hits/sample).

Claude Composite 41.2%
1X2
63.4% (225/355)
Total points
53.0% (141/266)
Exact score
0.4% (1/270)
Kimi Composite 39.9%
1X2
63.6% (124/195)
Total points
51.5% (84/163)
Exact score
0.0% (0/163)
Google AI Composite 37.7%
1X2
63.7% (283/444)
Total points
48.1% (199/414)
Exact score
0.0% (0/420)
ChatGPT Composite 36.6%
1X2
60.5% (46/76)
Total points
42.4% (25/59)
Exact score
0.0% (0/59)
GLM 5.2 Composite 35.8%
1X2
63.8% (143/224)
Total points
42.5% (94/221)
Exact score
0.9% (2/222)
Qwen Composite 34.6%
1X2
64.0% (126/197)
Total points
35.8% (63/176)
Exact score
0.6% (1/176)
DeepSeek Composite 29.6%
1X2
51.4% (19/37)
Total points
38.2% (13/34)
Exact score
0.0% (0/37)

The modèle's favorite by 1X2 = the max of P1/X/P2 in its probabilities; for sports without a nul (tennis, volleyball, etc.) the «X» option doesn't participate. grey — sample <5, not representative. Not betting advice.