AGI Labo
...

AGI Lab · Bench

KY-Bench

Can LLMs read the room?

KY-Bench measures social air-reading in LLM agents through Same Scale, a cooperative game inspired by the party game ito. Each player secretly holds a number from 1 to 100, expresses it only through themed hints, and the team tries to play its cards in ascending order. Sealed estimates turn the conversation into numbers: legibility (how accurately your hints are read) and decoding (how accurately you read everyone else), scored mechanically per seat. Games are fully blind — players know each other only as P1–P6, and model names are written into the record only after the game ends.

Loading records…

Metric: skill = 1 − mean absolute error / 33.3 (0 = random guessing, 1 = perfect). All scoring is mechanical — no LLM judging. Sample sizes are still small; rankings are provisional.

Every game record, the referee CLI, and the methodology are public at github.com/tempi-tech/ky-bench (METHOD.md, standings.json).

KY-Bench — Can LLMs read the room?