Anthropic used Pokémon to benchmark its newest AI model

Anthropic used Pokémon to benchmark its newest AI model. Yes, really.

In a blog post published Monday, Anthropic said that it tested its latest model, Claude 3.7 Sonnet, on the Game Boy classic Pokémon Red. The company equipped the model with basic memory, screen pixel input, and function calls to press buttons and navigate around the screen, allowing it to play Pokémon continuously.

A unique feature of Claude 3.7 Sonnet is its ability to engage in “extended thinking.” Like OpenAI’s o3-mini and DeepSeek’s R1, Claude 3.7 Sonnet can “reason” through challenging problems by applying more computing — and taking more time.

That came in handy in Pokémon Red, apparently.

Compared to a previous version of Claude, Claude 3.0 Sonnet, which failed to leave the house in Pallet Town where the story begins, Claude 3.7 Sonnet successfully battled three Pokémon gym leaders and won their badges.

Image Credits:Anthropic

Now, it’s not clear how much computing was required for Claude 3.7 Sonnet to reach those milestones — and how long each took. Anthropic only said that the model performed 35,000 actions to reach the last gym leader, Surge.

It surely won’t be long before some enterprising developer finds out.

Pokémon Red is more of a toy benchmark than anything. However, there is a long history of games being used for AI benchmarking purposes. In the past few months alone, a number of new apps and platforms have cropped up to test models’ game-playing abilities on titles ranging from Street Fighter to Pictionary.

Source link

Anthropic used Pokémon to benchmark its newest AI model

Recent posts

Lingo.dev is an app localization engine for developers

DoorDash’s app can now import your grocery list for faster shopping

Sonos delays set-top box after flawed app update

An Apple employee is suing the company over monitoring employee personal devices

Entrepreneur Marc Lore on ‘founder mode,’ bad hires, and why avoiding risk is deadly

WhatsApp is adding a way to turn selfies into stickers

Google CEO Sundar Pichai announces $120M fund for global AI education

Nvidia just became the world’s largest company amid AI boom

Y Combinator’s next Demo Day will include in-person seats for top VCs, Garry Tan says

Women in AI: Dr. Rebecca Portnoff is protecting children from harmful deepfakes

Three new ways to personalize your iPhone’s Home Screen in iOS 18

Spotify debuts marketing tools and insights for audiobook authors

Edera is building a better Kubernetes and AI security solution from the ground up

Cohere co-founder Nick Frosst’s indie band, Good Kid, is almost as successful as his AI company

Instagram Reels adds new features as TikTok is banned in the US

Related articles

DOGE’s HR email is getting the ‘Bee Movie’ spam treatment

SpaceX says Starship self-destructed after propellant leaks caused fires and comms blackout

Grok 3 appears to be driving Grok usage to new heights

Perplexity teases a web browser called Comet

The startup battle begins now: TechCrunch Startup Battlefield 200 applications are open

Flexport releases onslaught of AI tools in a move inspired by ‘Founder Mode’

A single default password exposes access to dozens of apartment buildings

Meta AI arrives in the Middle East and Africa with support for Arabic

Company

Follow us