Hacker Newsnew | past | comments | ask | show | jobs | submitlogin
Claude, GPT, Gemini Play Doom II and 19 More (VideoGameBench) (vgbench.com)
4 points by foddiangames 4 months ago | hide | past | favorite | 1 comment


Researchers from Princeton introduce a research preview of VideoGameBench, a benchmark which challenges vision-language models to complete, in real-time, a suite of 20 different popular video games from both hand-held consoles and PC.

GPT-4o, Claude Sonnet 3.7, Gemini 2.5 Pro, and Gemini 2.0 Flash playing Doom II (default difficulty) on VideoGameBench-Lite with the same input prompt! Models achieve varying levels of success but none are able to pass even the first level.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: