Fireship
Level B2· 431 words
Analysed: 97 expressions
Last week, Chinese AI lab Moonshot dropped Kimi K3, a massive open-source monster that instantly parameter mogged every other open model in existence. And not just that, it has OpenAI and Anthropic terrified because its trust-me-bro benchmark performance is on par with, and in some cases beating,
Claude Fable and GPT-5.6 Soul. And that's crazy because just a few weeks ago, the government was telling you these models were too dangerous for the common man to use. It only took the Chinese a few weeks to match their performance and then give away all the weights for free. But big AI is not too
happy about this, and there are already calls to ban Chinese models completely in the United States. In today's video, we'll take a look at the massive breakthrough that is Kimi K3 and find out what it means for the future of artificial intelligence. It is July 22nd, [music] 2026, and you're watching The Code
Report. Like Claude Fable and GPT-5.6 Soul, Kimi K3 is a native multi-model mixture of experts model with a 1 million token context window and a staggering 2.8 trillion parameters, and is optimized for jobs like long-horizon reasoning and, of course, coding. One interesting characteristic of K3 as a
mixture of experts model is that it has 896 total experts, of which exactly 16 activate per token, which means it works just like a big corporation where 16 good programmers do all the work while 880 other managers sit there and do nothing. But apparently, this works out well for Kimi because it makes scaling
about 2.5 times more efficient than K2. Despite these gains in efficiency, Kimi was so popular upon release that their GPUs ran out of juice and they had to start turning away paying customers. And as of making this video, all the paid plans are currently sold out. That's unfortunate, but in theory, you could
self-host the model because the weights are open. The weights are expected to be released on July 27th, but there's no chance in hell you'll be able to run it on your little gaming GPU. To run a monster like this, you'll need a massive array of data center caliber GPUs. But in theory, if you could do that, you
would have unlimited access to a Fable Soul caliber model, and that would be amazing because the benchmark situation with K3 is pretty wild. K3 is ranked number one on front end code Arena at a 1,679 Elo, which puts it ahead of Fable 5 and GPT-5.6 Soul. In addition, it lands in the top three on the artificial analysis
intelligence index. And if we look at every other coding benchmark, it's at least very competitive with the other frontier models. But you should never trust the trust me bro benchmarks because many of the K3 numbers were produced with Moon Shot's own Kimiko harness while competitors ran in
different harnesses. That could make Kimiko look slightly better at coding, but to their credit, Moon Shot admits that K3 still trails Fable and GPT-5.6 Soul overall, especially on benchmarks like Humanity's Last Exam where it's down by about 10 points. On top of that, artificial analysis measured a 51%
hallucination rate, which is definitely not a good thing, especially when it comes to coding. In addition, it also tends to spit out way more tokens than it needs to, which could ultimately end up costing you more money despite the model itself being cheaper. When it comes to things like UI design and data
visualization, it's extremely impressive for an open model, but in my opinion, it's still one step behind Fable and GPT Soul. But one of the most interesting things about this release is the geopolitics surrounding it. Recently at the World AI conference, China's Communist Party became the loudest