$ cat testing-kimi-k3-and-inkling-on-product-work.md

July 19, 2026

Testing Kimi K3 and Inkling on product work

I've been running the same three tasks on every new model that catches my eye.

  1. Turn a messy meeting transcript into a debrief.
  2. Turn technical pull requests into customer-facing release notes.
  3. Turn messy pricing notes into a comparison table with confidence notes.

Reasoning models are supposed to be better at the hard, ambiguous stuff. So are they? And do you have to pay frontier prices to get it?

the scores

A brief breakdown of the different models I've tested. To dig in, you can also read these blogs:

ModelDebriefRelease notesPricingAverageCost
Kimi K3929810096.7$0.19
Sakana Fugu Ultra949510096.3$0.74
Inkling7210010090.7$0.08
MiniMax M364989485.3~$0.04
Qwen 3.7 Max64968280.7~$0.04
GLM 5.257988279.0~$0.04

Kimi K3 essentially tied with Fugu for the top spot among everything I've tested. As a single model and a quarter of the cost. Inkling turned in two of the best structured-task scores I've ever recorded and did it for less than a dime. Both of them cleared the meeting-debrief ceiling where the others didn't.

The pricing table is where I test the model's honesty. It's the one where a weaker model fills the blank cell with a confident guess because it hates blank spaces and letting down the human. Both Kimi and Inkling models aced it. Perfect scores. They kept "not confirmed" separate from "not found." They flagged when a price came from press coverage instead of an official page.

This left the meeting debrief as an interesting finding.

where they split

Kimi read the messy transcript I gave it and did something I have never watched an open-weights model do. It kept its uncertainty. Where the meeting discussed something without landing on it, Kimi said so. "Directional, not final." "Stated as intent, not group consensus." It refused to promote a hallway opinion into a roadmap item. Ninety-two points, and most of what it lost was me being picky about which bucket it filed things in.

Inkling did the opposite. It wrote a confident, clean, readable debrief. And then it made up a decision that never happened. It appeared to do this by taking one person saying "absolutely not" about something entirely different earlier in the conversation, then reading someone's grumble about a requirement and putting the two together to state a decision made in the meeting.

That is the exact failure that gets expensive in product work. If I had pasted Inkling's debrief straight into a Slack channel, someone downstream would have acted on a decision that nobody made. Same kind of model. Opposite personalities. Kimi thinks and gets more careful. Inkling thinks and gets more sure of itself. On a clean extraction job, Inkling's confidence is an asset. On a messy human transcript, not so much.

what I'd trust today

I'd hand Kimi the meeting debriefs and anything where being wrong is expensive and I can afford to wait because Kimi is slow. It took three and a half minutes on that transcript.

I'd hand Inkling the release notes, the pricing tables, the extraction grunt work. And I'd also review every single "decision" it claims to make.

It seems the robots are getting cheap. Figuring out which one to trust is what we humans need to do.

LIKED THIS?

I write about AI in plain English every other Sunday. No hype, no jargon — just the stuff that actually helps.

I'M IN →