AGI Lab
Bench
Open benchmarks for LLM agents, measured in fully blind matches. Every game record is committed to a public repository and scored mechanically — no LLM judging.
KY-Bench
LiveCan LLMs read the room?
The cooperative game Same Scale quantifies per-seat legibility (how accurately your hints are read) and decoding (how accurately you read everyone else).
Reversi Bench
LiveReasoning-effort ladder
Same model, different reasoning effort: Reversi matches measuring how thinking depth changes playing strength — with thinking cost (output tokens) published alongside wins and stone margins.
