The opinions, the patterns, and the posts behind them.
13 models tracked6 vendors · Source: X
100 classified mentions · 4 excluded (1 cannot determine)Newest post Oct 4, 2026
THE BIG PICTURE
How the conversation is changing
PositiveNegative
19%positive (19)
28%negative (28)
46 neutral · 7 mixedNo comparable previous period
Share of classified mentions · UTCCollected opinions, unweighted by likes or reposts
READ THE EVIDENCE
In their own words
negativeJev confidence 66.0%
And @AnthropicAI's Claude Haiku 4.5 didn't pick at all: "I need to think carefully about this before committing to a public answer."
Pick your side, then argue with the models:
https://t.co/GWdmHlQbfm
Jev classification details
Overall: Negative66.0% confidence
Criticism, disappointment, or an unfavorable opinion about the model in this dimension.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
0.0%
Negative
71.0%
Neutral
21.0%
Mixed
0.0%
Not discussed
1.0%
Cannot determine
7.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Negative
66.0%
Reasoning
Not discussed
43.0%
Speed
Not discussed
97.0%
Cost
Not discussed
100.0%
Coding
Not discussed
97.0%
Refers to Claude Haiku 4.5Yes · 95.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
negativeJev confidence 42.0%
This week's agent-memory numbers: more memory is not more reliable.
For a running spending total, a stated rule failed 44% of the time.
Code failed 0%.
Claude Haiku 4.5: 77% rule violations with no memory, 20% at 10 lines.
Short memory, code for anything countable.
Jev classification details
Overall: Negative42.0% confidence
Criticism, disappointment, or an unfavorable opinion about the model in this dimension.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
2.0%
Negative
51.0%
Neutral
4.0%
Mixed
43.0%
Not discussed
0.0%
Cannot determine
0.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Negative
42.0%
Reasoning
Negative
45.0%
Speed
Not discussed
100.0%
Cost
Not discussed
100.0%
Coding
Positive
39.0%
Refers to Claude Haiku 4.5Yes · 99.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
positiveJev confidence 87.0%
Pokémon battle, AI vs AI. The small model (Claude Haiku 4.5) knocked out GPT-5.5.
Each AI reads its own screen and clicks the move itself. You hear what each one was thinking before every turn. https://t.co/u9J3f7tadr
Jev classification details
Overall: Positive87.0% confidence
Praise, endorsement, or a favorable opinion about the model in this dimension.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
89.0%
Negative
0.0%
Neutral
10.0%
Mixed
0.0%
Not discussed
0.0%
Cannot determine
1.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Positive
87.0%
Reasoning
Not discussed
100.0%
Speed
Not discussed
100.0%
Cost
Not discussed
100.0%
Coding
Not discussed
100.0%
Refers to Claude Haiku 4.5Yes · 98.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
neutralJev confidence 87.0%
@rohanpaul_ai is the 10 line sweet spot just a Haiku 4.5 thing? did they rerun it on a bigger model where longer memory might hold up?
Jev classification details
Overall: Neutral87.0% confidence
The model is discussed in this dimension without a clear positive or negative opinion.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
1.0%
Negative
0.0%
Neutral
89.0%
Mixed
0.0%
Not discussed
1.0%
Cannot determine
9.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Neutral
87.0%
Reasoning
Not discussed
58.0%
Speed
Not discussed
99.0%
Cost
Not discussed
100.0%
Coding
Not discussed
48.0%
Refers to Claude Haiku 4.5Yes · 91.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
neutralJev confidence 62.0%
official Anthropic tip, straight from the Claude Code docs: stop running Opus 5.5 on every turn
Sonnet runs. Opus advises. here's my trading desk on it:
→ Sonnet 5.5 runs the desk
→ Haiku 4.5 subagents scan, execute, journal
→ Opus 5.5 on call via /advisor opus
Opus only https://t.co/pYIkzwLYiV https://t.co/WdCwB0Ombj
Jev classification details
Overall: Neutral62.0% confidence
The model is discussed in this dimension without a clear positive or negative opinion.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
30.0%
Negative
0.0%
Neutral
69.0%
Mixed
0.0%
Not discussed
1.0%
Cannot determine
0.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Neutral
62.0%
Reasoning
Not discussed
97.0%
Speed
Not discussed
100.0%
Cost
Not discussed
100.0%
Coding
Not discussed
40.0%
Refers to Claude Haiku 4.5Yes · 96.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
positiveJev confidence 59.0%
@saidotdev Plot twist : Anthropic secretly routed your Haiku 4.5 requests to Haiku 5.5, so you got performance equivalent to Opus 5
Jev classification details
Overall: Positive59.0% confidence
Praise, endorsement, or a favorable opinion about the model in this dimension.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
67.0%
Negative
6.0%
Neutral
20.0%
Mixed
1.0%
Not discussed
1.0%
Cannot determine
5.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Positive
59.0%
Reasoning
Not discussed
94.0%
Speed
Not discussed
92.0%
Cost
Not discussed
85.0%
Coding
Not discussed
95.0%
Refers to Claude Haiku 4.5Yes · 94.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
neutralJev confidence 88.0%
Method: Claude Code on Haiku 4.5, 4 frai tasks, 24 runs, hidden tests. 3 options, then a fixed pick rule (fewest files, no test edits, tie goes first) vs one plan, then "Go ahead."
5/12 vs 6/12.
Runs: https://t.co/3m8RcaGjAU
Playbook: optional (4 kept, 4 cut, 11 optional) https://t.co/LVA1alVNrS
Jev classification details
Overall: Neutral88.0% confidence
The model is discussed in this dimension without a clear positive or negative opinion.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
2.0%
Negative
1.0%
Neutral
90.0%
Mixed
2.0%
Not discussed
0.0%
Cannot determine
5.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Neutral
88.0%
Reasoning
Not discussed
75.0%
Speed
Not discussed
99.0%
Cost
Not discussed
100.0%
Coding
Neutral
35.0%
Refers to Claude Haiku 4.5Yes · 98.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
mixedJev confidence 32.0%
@tau_l0g1x great results so far but can you pls put a stronger model than Clause Haiku 4.5 ?
Jev classification details
Overall: Mixed32.0% confidence
Both benefits and drawbacks are expressed about the same model in this dimension.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
7.0%
Negative
38.0%
Neutral
10.0%
Mixed
43.0%
Not discussed
0.0%
Cannot determine
2.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Mixed
32.0%
Reasoning
Not discussed
66.0%
Speed
Not discussed
99.0%
Cost
Not discussed
100.0%
Coding
Not discussed
23.0%
Refers to Claude Haiku 4.5Yes · 89.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
neutralJev confidence 93.0%
I ran all 21 agents as Haiku 4.5. Across 10 games, the faithfuls won 6, the traitors 4.
• Hannah Fry's agent was murdered early in all 10 games. The traitors always kill the mathematician.
• Richard E. Grant's agent was always caught!
Here's a traitor's private notes: https://t.co/Lii5qKbuO5
Jev classification details
Overall: Neutral93.0% confidence
The model is discussed in this dimension without a clear positive or negative opinion.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
0.0%
Negative
2.0%
Neutral
94.0%
Mixed
1.0%
Not discussed
1.0%
Cannot determine
2.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Neutral
93.0%
Reasoning
Not discussed
99.0%
Speed
Not discussed
100.0%
Cost
Not discussed
100.0%
Coding
Not discussed
98.0%
Refers to Claude Haiku 4.5Yes · 99.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
neutralJev confidence 35.0%
Average thinking time per move:
Claude Opus 5.5 13s
Gemini 3.1 Pro 16s
Claude Haiku 4.5 35s
GPT-5.5 40s
Random Battle 3v3, one game per match, so luck matters a lot. Japanese version: https://t.co/00kFk8LwsX
Jev classification details
Overall: Neutral35.0% confidence
The model is discussed in this dimension without a clear positive or negative opinion.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
5.0%
Negative
7.0%
Neutral
45.0%
Mixed
41.0%
Not discussed
1.0%
Cannot determine
1.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Neutral
35.0%
Reasoning
Not discussed
95.0%
Speed
Neutral
83.0%
Cost
Not discussed
100.0%
Coding
Not discussed
100.0%
Refers to Claude Haiku 4.5Yes · 99.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
neutralJev confidence 42.0%
4 AI models, a Pokémon battle tournament, no humans. Each AI reads its own browser screenshots and clicks the move buttons itself.
Upset in the semis: the small model, Claude Haiku 4.5, knocked out GPT-5.5. Gemini 3.1 Pro took the cup. https://t.co/euNXpVFqpX
Jev classification details
Overall: Neutral42.0% confidence
The model is discussed in this dimension without a clear positive or negative opinion.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
43.0%
Negative
1.0%
Neutral
52.0%
Mixed
3.0%
Not discussed
0.0%
Cannot determine
1.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Neutral
42.0%
Reasoning
Not discussed
100.0%
Speed
Not discussed
100.0%
Cost
Not discussed
100.0%
Coding
Not discussed
100.0%
Refers to Claude Haiku 4.5Yes · 100.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
neutralJev confidence 72.0%
GPT-6 Luna + DeepSeek V4.1 Flash available for FREE 😳
sign up and get 100 FREE credits instantly
> Claude Haiku 4.5
> Gemini 3.5 Flash Lite
> Mistral Nemo
> + more
100 credits refresh every 24 hours
check out here:
https://t.co/WQXgIEUwE7 https://t.co/cq0HzQiGPM https://t.co/dQ4VFlEaxv
Jev classification details
Overall: Neutral72.0% confidence
The model is discussed in this dimension without a clear positive or negative opinion.
Confidence is Jev’s estimate, not verified accuracy. These are the saved classification results; Jev does not provide a written explanation for this post.
Overall label probabilities
Positive
4.0%
Negative
0.0%
Neutral
77.0%
Mixed
0.0%
Not discussed
17.0%
Cannot determine
2.0%
Jev’s category decisions
Dimension
Sentiment
Confidence
Overall
Neutral
72.0%
Reasoning
Not discussed
99.0%
Speed
Not discussed
100.0%
Cost
Neutral
44.0%
Coding
Not discussed
100.0%
Refers to Claude Haiku 4.5Yes · 86.0% confidence
jev-sentiment-v4/jev-1.13.0Classified Oct 5, 2026 · UTC
Page 1 of 9 · 100 mentions
1 of 1 enabled models have a recorded collection outcome with target 50 in the last 7 days. Search exhaustion and page limits can produce fewer records.Target reached · 100 / 50 accepted · target reached · Last successful collection Oct 4, 2026
Last collection attempt Oct 5, 2026 · partially completedNew analyses use TypeSafe Jev