Back to Home

Qwen 3.5 Small: Tiny Models, Massive Impact?

Qwen 3.5 Small: The Specs Worth Squinting At

Alibaba dropped its Qwen 3.5 Small Model series on a Monday — four sizes: 0.8B, 2B, 4B, and 9B parameters, each one a native multimodal that can see images, process text, and (on the bigger two) reason its way through a problem without melting your GPU. By Tuesday, the project's technical lead was gone. That timing is either a coincidence or a signal, and in the AI world, coincidences this juicy rarely are.

Let's get the numbers out first, because the models actually deserve the attention.

Qwen 3.5 Small vs. The Competition: Head-to-Head

The small-model arena in 2026 is ruthlessly crowded. Meta's Llama 3.2 (1B and 3B), Microsoft's Phi-3.5, Google's Gemma 2 (2B and 9B), and now Qwen 3.5 Small are all fighting for the same slice of the pie: on-device inference, lightweight agents, and budget-friendly fine-tuning. Here's how they stack up:

  • Parameter Efficiency: Qwen 3.5's 9B variant punches above its weight class. Early community benchmarks show it matching Llama 3.2 3B on MMLU-Pro while being 3x larger in parameter count — but that's not the full story. Qwen packs native vision into every single size, including the 0.8B model. Neither Llama 3.2 nor Gemma 2 does that at their smallest tiers. Phi-3.5 vision starts at 4.2B. Advantage: Qwen.
  • Inference Speed: On an M4 MacBook Air, the 2B Qwen 3.5 Small clocks ~45 tokens/second with 4-bit quantization. The comparable Gemma 2 2B hits ~52 tokens/second but lacks vision support. Llama 3.2 3B (text-only) manages ~40 tokens/second. If you need a multimodal model running at edge latency, Qwen 3.5 Small is the current leader.
  • License and Friction: All Qwen 3.5 Small models are open-weight under the Qwen License, which permits commercial use but includes restrictions on competitive benchmarking and requires attribution. Llama 3.2 uses the Llama 3.2 Community License (more permissive). Gemma 2 uses the Gemma Terms (broadly permissive). If you're building a product that publishes benchmarks, Qwen's license is the most restrictive of the three.

The Elephant in the Server Room: Junyang Lin's Exit

Junyang Lin, the central technical leader who had been with Qwen since April 2023, posted on X that he was "stepping down" the day after the 3.5 Small launch. His LinkedIn still showed him at Alibaba — until it didn't. Colleagues reacted in ways that suggest this wasn't a planned handoff. One called it "the end of an era." Another wrote directly to Lin: "I know leaving wasn't your choice." A second Qwen contributor, Binyuan Hui, quietly updated his X bio to "formerly MTS @Alibaba_Qwen."

Alibaba declined to comment. Lin declined to comment. The silence is louder than anything either could have said.

Elon Musk, meanwhile, was busy praising the models on X — "impressive intelligence density" — apparently unaware or unconcerned about the internal drama unfolding around the team he was complimenting.

Before vs. After Lin: What Changes?

Comparing Qwen's trajectory with Lin at the helm versus after his departure is a thought exercise, but a revealing one.

  • Release Cadence: Under Lin, Qwen shipped at a relentless pace — Qwen (Apr 2023), Qwen 2 (mid-2024), Qwen 2.5 (late 2024), Qwen 3 (Apr 2025), and now Qwen 3.5 (Mar 2026). That's roughly one major release every 6-8 months. Whether a new lead can sustain that rhythm is the open question.
  • Community Engagement: Lin was the public face of Qwen on X, engaging with developers, absorbing feedback, and building the global bridges that turned Qwen from an Alibaba internal project into an internationally recognized open-weight contender. His replacement will need to earn that trust from scratch.
  • Technical Direction: The Qwen 3.5 Small series leans hard into multimodal capability at every scale — a strategy that differentiates it from Llama and Gemma. Will the next leadership double down or pivot?

The Verdict

On technical merit alone, Qwen 3.5 Small is the best open-weight small-model family available today if you need native multimodality across all sizes. The 9B variant holds its own against models twice its size, the 0.8B model running on a phone can describe images in real time, and the community tooling (llama.cpp, Ollama, MLX) already supports it. For text-only use cases, Llama 3.2 and Gemma 2 remain competitive and carry fewer licensing headaches.

On organizational stability, the score is worrying. Losing a technical lead mid-launch, with hints of involuntary departure and a cascade of team exits, does not inspire confidence in Qwen's 2026-2027 roadmap. Alibaba's commitment to the project is not in doubt — the investment is too large and the geopolitical stakes too high. But software engineering runs on people, and Qwen just lost a foundational one.

Winner (Technical): Qwen 3.5 Small
Winner (Long-Term Confidence): TBD — watch the next 90 days for a new tech lead announcement and the next release cadence.

If you're building a multimodal agent today and need something that runs on-device, Qwen 3.5 Small is the play. If you're planning a deployment six months out, you might want to keep one eye on the talent pipeline.

Comments

No comments yet. Be the first to share your thoughts!