Open vs. Closed: Did the Frontier Labs Bet on the Wrong Horse? A Kimi K3 Case Study
Kimi K3 launched July 16. On July 17 I tested it against five models, including Claude Fable 5 and GPT-5.6 Sol. It won. First place on overall preference, 2.2x the second place model. It also fabricated a statistic and failed my factual gate on half its outputs. Every K3 benchmark you have seen measures coding. […]
