Kimi K3 launched July 16. On July 17 I tested it against five models, including Claude Fable 5 and GPT-5.6 Sol.
It won. First place on overall preference, 2.2x the second place model.
It also fabricated a statistic and failed my factual gate on half its outputs.
Every K3 benchmark you have seen measures coding. I tested it as a technical writer: read a research paper, write an audience-ready brief, and get every number right. That is the job your marketing, training, and sales ops teams do every day.
Three findings from 240 judged comparisons:

Open vs. closed predicted nothing. The two perfect factual records were one closed model and one open model. The worst record was a closed model.
Style beat substance. The AI judges preferred the model that made up a number over the model with a perfect record. Asked to score accuracy alone, they reversed and put Sol first.
Judges cannot agree on truth. My two independent judges agreed on factual fidelity only 20% of the time. On writing quality, 58%.
If your AI content evaluation ends at “the judge preferred it,” you have a style contest, not a quality gate. Verify the numbers against the source first. That deterministic check was the only signal in my pipeline that agreed with itself.
Full paper in the comments.
#IncentiveCompensation #SalesOps #AIEvaluation #KimiK3
