Back to Home

Qwen's 'Free' 2.4T Flagship vs the 27B That Ate Hacker News

Alibaba just had the strangest week in open-source AI since Kimi K3 taught the world that "open" can come with a calculator attached. The Qwen family dropped a double feature: a 2.4-trillion-parameter flagship whose free weights demand a rack of GPUs that costs more than most startups, and a humble 27B model that runs on a gaming PC and became the most-loved thing on Hacker News all week.

One of these models made headlines. The other made developers happy. They are not the same model.

The Flagship: Free Weights, Fine Print Not Included

On August 12, Alibaba finally shipped the open weights for Qwen3.8-Max: the Qwen3.8-2.4T-A95B checkpoint (plus an FP8 variant) landed on Hugging Face, and NVIDIA's own deployment blog confirmed it the same day, while casually noting that serving it properly means a GB300 NVL72 rack. That is 72 Blackwell Ultra GPUs of "properly."

Here is the thing about free weights: the BF16 checkpoint is 4.89 terabytes. Even aggressively quantized, you are looking at 397 GB to 1.31 TB. Your gaming rig is not going to cut it. Your workstation is not going to cut it. Your entire office might not cut it.

And the license is not the Apache 2.0 that every earlier Qwen generation shipped under. It is a bespoke qwen3.8-max license with a reported revenue-share clause: build a model-as-a-service or AI assistant business on top of it, cross $50 million in annual revenue, and Alibaba wants a conversation. A $50 million safe harbor for startups sounds generous, right up until you notice it is calibrated to catch exactly the companies that could threaten Alibaba's own API business. This is not a gift. It is a toll booth.

The community clocked it fast. Within hours, the Hugging Face discussion thread was calling the stripped-down release a "huge disappointment." No vision. No 1M-token context. Text-only and thinking-mode-only, wearing the same name as the API model that has all the good parts. One commenter compared it to game-industry DLC paywalling. Another said it cost Alibaba "all good faith" built up over years of Qwen releases. Ouch.

Early traction? A few hundred likes and under a thousand downloads for the flagship of the world's most-downloaded open model family. Hacker News was not impressed.

The Underdog: 27 Billion Parameters of Community Gold

Three days later, on August 15, the other half of the promise shipped. Qwen3.8-27B is a dense 27B vision-language model with 262K native context and an Apache 2.0 license. The real Apache, the good one. It runs on a single RTX 4090. It runs on a Mac Studio. Someone on Hacker News got it running on a DGX Spark.

Hacker News put it at #1 with 893 points, one of the biggest AI threads of the week. Unsloth had GGUF quantizations up within hours, and the local-AI crowd went from calling it a "huge disappointment" to hailing a "renewal of the beloved Qwen model" in about 72 hours.

About Those "Beats Claude Opus" Headlines

Alibaba's own model card puts Qwen3.8-27B within striking distance of Claude Opus-class scores on agentic coding. On SWE-bench Pro it edges the cited Opus figure, 61.7 vs 53.4. That single line is doing most of the "beats Opus" work in the headlines.

The honest read: it wins one of the five benchmarks in Alibaba's own comparison and trails on the other four, sometimes by a lot (HLE: 30.8 vs 40.0). Different harnesses, different temperatures, and a recurring "benchmaxxing" suspicion make apples-to-apples claims shaky. Genuinely impressive for a 27B dense model. Just hold the "beats Opus" claim loosely.

So where does that leave us?

  • The 2.4T flagship: open weights, closed revenue ceiling, and a hardware bill that reads like a mortgage.
  • The 27B: genuinely free, genuinely local, genuinely loved. The release that actually matters.
  • The play: the small model captures developers, the big model gates the enterprise. Two-pronged, no apology.

Alibaba is positioning the 27B as proof that open weights now lag frontier closed models by about six months instead of by forever. Whether that survives independent testing is a question for another week. For now, the internet has voted, and the little guy won.

One tip before you download: leave reasoning_effort on medium. The xhigh setting is powerful, but it has been caught spending 20 to 90 minutes on a single SVG. That is not thinking. That is rumination.

Comments

No comments yet. Be the first to share your thoughts!