devwithjev
— reading now— views
Submit a build

Jev Plays Snake Against Six LLMs

Jev played Snake against six LLMs using the same seed and four-way decision space for 60 seconds each. It scored 21 points across 204 decisions, the highest score in the experiment.

Angel Galvis Caballero@angelgalvisc𝕏
“Models have been superhuman at chat for years. So where is all the automation?” I took @CompleteSkeptic’s question literally and built a small experiment. Jev vs. six LLMs playing Snake. Same seed, same four-way decision space, 60 seconds each. Jev: 21 pts / 204 decisions Haiku: 9 / 73 Luna: 5 / 43 Sol: 5 / 40 Kimi K2.7: 5 / 34 Opus: 3 / 25 Kimi K3: 2 / 14 My first thought was: maybe Jev is simpl
Sep 20, 2026X postsView on X
Angel Galvis Caballero built the experiment in response to @CompleteSkeptic’s question about where automation is. He described it as a small experiment.

Also filed under Games & real time

  • Benchmarks Jev on Pokémon Red

    The project benchmarks Jev, TypeSafe’s fast decision model, against Jev paired with a GPT-6 Sol planner for playing Pokémon Red.

  • Plain-English Cellular Ecosystem Simulator

    LifePot is a cellular automaton-inspired ecosystem that users describe in plain English. Jev turns those inputs into species, feeding relationships, reproduction strategies, and environmental conditions.

  • Mapping Children's Behavioral Intent in Games

    Mores Research used Jev to map children’s behavioral intent in games and examine how it correlates with real-world behavior. It reports mapping intents across 45+ real sessions involving 20 child users over a long horizon.

  • StarCraft Agent for Fast Tactical Decisions

    Jev_Star is a StarCraft agent experiment that uses Jev for fast tactical decisions.