People are using Super Mario to benchmark AI now

Date:

Share post:


Thought Pokémon was a tough benchmark for AI? One group of researchers argues that Super Mario Bros. is even tougher.

Hao AI Lab, a research org at the University of California San Diego, on Friday threw AI into live Super Mario Bros. games. Anthropic’s Claude 3.7 performed the best, followed by Claude 3.5. Google’s Gemini 1.5 Pro and OpenAI’s GPT-4o struggled.

It wasn’t quite the same version of Super Mario Bros. as the original 1985 release, to be clear. The game ran in an emulator and integrated with a framework, GamingAgent, to give the AIs control over Mario.

Image Credits:Hao Lab

GamingAgent, which Hao developed in-house, fed the AI basic instructions, like, “If an obstacle or enemy is near, move/jump left to dodge” and in-game screenshots. The AI then generated inputs in the form of Python code to control Mario.

Still, Hao says that the game forced each model to “learn” to plan complex maneuvers and develop gameplay strategies. Interestingly, the lab found that reasoning models like OpenAI’s o1, which “think” through problems step by step to arrive at solutions, performed worse than “non-reasoning” models, despite being generally stronger on most benchmarks.

One of the main reasons reasoning models have trouble playing real-time games like this is that they take a while — seconds, usually — to decide on actions, according to the researchers. In Super Mario Bros., timing is everything. A second can mean the difference between a jump safely cleared and a plummet to your death.

Games have been used to benchmark AI for decades. But some experts have questioned the wisdom of drawing connections between AI’s gaming skills and technological advancement. Unlike the real world, games tend to be abstract and relatively simple, and they provide a theoretically infinite amount of data to train AI.

The recent flashy gaming benchmarks point to what Andrej Karpathy, a research scientist and founding member at OpenAI, called an “evaluation crisis.”

“I don’t really know what [AI] metrics to look at right now,” he wrote in a post on X. “TLDR my reaction is I don’t really know how good these models are right now.”

At least we can watch AI play Mario.



Source link

Lisa Holden
Lisa Holden
Lisa Holden is a news writer for LinkDaddy News. She writes health, sport, tech, and more. Some of her favorite topics include the latest trends in fitness and wellness, the best ways to use technology to improve your life, and the latest developments in medical research.

Recent posts

Related articles

Moonwatt secures $8.3M to dial up solar’s staying power with sodium-ion storage

The drive to decarbonize our economies through electrification and clean energy continues to generate momentum around battery...

Trump Administration cuts may threaten AI research efforts

The Trump Administration has fired a number of National Science Foundation employees who had been handpicked for...

General Catalyst loses three top investors as the firm expands beyond venture, contemplates IPO

Three key investors have left General Catalyst amidst a series of recent changes at the firm, which...

You can now talk to Google Gemini from your iPhone’s lock screen

Google Gemini users can now access the AI chatbot directly from the iPhone’s lock screen, thanks to...

Singapore arrests alleged Nvidia chip smugglers

There’s lots of scrutiny involving China obtaining advanced Nvidia chips despite strict U.S. export controls, with Chinese...

Creator monetization platform Passes sued over alleged distribution of CSAM

Passes, a direct-to-fan monetization platform for creators backed by $40 million in Series A funding, has been...

MWC hears two starkly divided views of AI’s impact

Two sharply different visions of AI were platformed on stage at the Mobile World Congress trade show...

The author of SB 1047 introduces a new AI bill in California

The author of California’s SB 1047, the nation’s most controversial AI safety bill of 2024, is back...