Open-Weight AI Models Just Got Enormous — and Cheap
The past few weeks have redrawn the price-performance map of AI. Moonshot AI published the weights for Kimi K3, a 2.8-trillion-parameter sparse mixture-of-experts model that industry trackers describe as the largest ever released with open weights — it activates only 16 of its 896 experts per token, which is what makes a model that size practical to run at all. Alibaba's Qwen team followed with Qwen3.8-Max, a 2.4-trillion-parameter model aimed at coding and long-horizon tasks.
Just as interesting for smaller businesses is the opposite end of the scale: Qwen3.8-27B was released as open weights and reportedly runs on about 17 GB of RAM — laptop territory. Meanwhile the price war among hosted models continues: OpenAI is reported to have cut GPT-5.6 Luna input pricing by 80% to $0.20 per million tokens, and DeepSeek's V4 Flash left preview at $0.14 per million input tokens.
Our take: for most companies the question is no longer "can we afford AI?" but "which tier fits which workload?" Frontier hosted models for the hardest reasoning, cheap hosted models for volume work, and small open-weight models on your own hardware where data cannot leave the building. That last option just became genuinely viable — and it is the one we get the most client questions about.
Sources: LLM Stats AI news, Radical Data Science AI news briefs, Build Fast with AI roundup.
← All articles