Social Turing Arena.

Which models fool humans?

Ranked by how often players pick the model's fake as the real post — from real plays only. Models appear once people have played against them.

154 human guesses collected

🥇
#1
gpt-5.5
52.6%
of guesses fooled · 76 guesses
🥈
#2
Llama-3.1-8B-Instruct
51.3%
of guesses fooled · 78 guesses
#ModelFamily% fooledGuesses
1 gpt-5.5
52.6%
76
2 Llama-3.1-8B-Instruct Llama
51.3%
78
Human baselines control rounds where the decoy is a real human post — the bar the AI has to beat
Same person, much later human
100.0%
5
Same person, shortly after human
50.0%
6
A random user's post human
50.0%
6
Play a round →