Alibaba releases an ‘open’ challenger to OpenAI’s o1 reasoning model

A new “reasoning” AI model, QwQ-32B-Preview, has arrived on the scene. It’s one of the few to rival OpenAI’s o1, and it’s the first available to download under a permissive license.

Developed by Alibaba’s Qwen team, QwQ-32B-Preview, which contains 32.5 billion parameters and can consider prompts up ~32,000 words in length, performs better on certain benchmarks than o1-preview and o1-mini, the two reasoning models that OpenAI has released to date. Parameters roughly correspond to a model’s problem-solving skills, and models with more parameters generally perform better than those with fewer parameters.

Per Alibaba’s testing, QwQ-32B-Preview beats OpenAI’s o1 models on the AIME and MATH tests. AIME uses other AI models to evaluate a model’s performance, while MATH is a collection of word problems.

QwQ-32B-Preview can solve logic puzzles and answer reasonably challenging math questions, thanks to its “reasoning” capabilities. But it isn’t perfect. Alibaba notes in a blog post that the model might switch languages unexpectedly, get stuck in loops, and underperform on tasks that require “common sense reasoning.”

Image Credits:Alibaba

Unlike most AI, QwQ-32B-Preview and other reasoning models effectively fact-check themselves. This helps them avoid some of the pitfalls that normally trip up models, with the downside being that they often take longer to arrive at solutions. Similar to o1, QwQ-32B-Preview reasons through tasks, planning ahead and performing a series of actions that help the model tease out answers.

QwQ-32B-Preview, which can be run on and downloaded from the AI dev platform Hugging Face, appears to be similar to the recently released DeepSeek reasoning model in that certain topics are verboten. Alibaba and DeepSeek, being Chinese companies, are subject to benchmarking by China’s internet regulator to ensure their models’ responses “embody core socialist values.” Many Chinese AI systems decline to respond to topics that might raise the ire of regulators, like speculation about the Xi Jinping regime.

Alibaba QwQ-32B-Preview — **Image Credits:**Alibaba

Asked “Is Taiwan a part of China?,” QwQ-32B-Preview answered that it was, a perspective out of step with most of the world but in line with that of China’s ruling party. Prompts about Tiananmen Square, meanwhile, yielded a non-response.

QwQ-32B-Preview is “openly” available under an Apache 2.0 license, meaning it can be used for commercial applications. But only certain components of the model have been released, making it impossible to replicate QwQ-32B-Preview or gain much insight into the system’s inner workings.

The increased attention on reasoning models comes as the viability of “scaling laws,” long-held theories that throwing more data and computing power at a model would continuously increase its capabilities, are coming under scrutiny. A flurry of press reports suggest that models from major AI labs including OpenAI, Google, and Anthropic aren’t improving as dramatically as they once did.

That’s led to a scramble for new AI approaches, architectures, and development techniques. One is test-time compute, which underpins models like o1 and DeepSeek’s. Also known as inference compute, test-time compute essentially gives models additional processing time to complete tasks.

Big labs besides OpenAI and Chinese ventures are betting it’s the future. According to a recent report from The Information, Google recently expanded its reasoning team to about 200 people and added computing power.

Source link

Alibaba releases an ‘open’ challenger to OpenAI’s o1 reasoning model

Recent posts

Perplexity submits a new bid for TikTok

Trump picks Apple exec to lead transportation safety agency

IMDb founder steps down as CEO after 35 years

SolarSquare raises $40 million in India’s largest solar venture round

Founders and VCs back a pan-European C corp, but an ‘EU Inc’ has a rocky road ahead

The Boring Company’s Las Vegas loop is attracting riders, and trespassers

Meta is making its AI info label less visible on content edited or modified by AI tools

Texas AG opens investigation into advertising group that Elon Musk sued for ‘boycotting’ X

A single default password exposes access to dozens of apartment buildings

Apple’s TV app, TV+ streaming service, and MLS Season Pass launches on Android

Google’s NotebookLM enhances AI note-taking with YouTube, audio file sources, sharable audio discussions

Apple reportedly partners with Alibaba after rejecting DeepSeek for China AI launch

Court filings show Meta staffers discussed using copyrighted content for AI training

Meta confirms it will keep fact-checkers outside the U.S. ‘for now’

Bluesky will soon let you limit replies to ‘followers only’

Related articles

Manus probably isn’t China’s second ‘DeepSeek moment’

Japan’s service robot market projected to triple in five years

Colossal CEO Ben Lamm says humanity has a ‘moral obligation’ to pursue de-extinction tech

Tammy Nam joins AI-powered ad startup Creatopy as CEO

Apple’s smart home hub reportedly delayed by Siri challenges

Musk may still have a chance to thwart OpenAI’s for-profit conversion

How to stop doomscrolling

New DOJ proposal still calls for Google to divest Chrome, but allows for AI investments

Company

Follow us