Home

/

Kestrel Labs’ open-weights model just beat the incumbents at code review. The incumbents are not amused.

The 40-person Oslo startup released K2-Review under an Apache license on Tuesday. Early benchmarks put it ahead of three closed models on real-world pull requests, at a tenth of the cost.

Xin

On Tuesday morning, a 40-person company in Oslo did something its far larger rivals have spent a year saying was too risky. Kestrel Labs published the full weights of K2-Review, a model trained only to read code changes and comment on them, under the Apache 2.0 license.

By the afternoon, three independent teams had run it against their own review queues. All three told Modelodge it caught more real bugs than the closed models they pay for today.

What K2-Review actually does

K2-Review is small by current standards: 34 billion parameters, fine-tuned on 11 million annotated pull requests. It does not write code. It reads a diff, the surrounding files and the ticket, then leaves line-level comments the way a senior engineer would.

The benchmark that matters

Kestrel’s headline number comes from a test set of 4,200 merged pull requests that were later reverted because they caused an incident. K2-Review flagged 71% of them. The best closed model in the comparison flagged 58%.

Nobody gets promoted for the bug that didn’t ship. That’s why review never got the attention generation did.

Ingrid Solberg · CTO, Kestrel Labs

What happens next

Kestrel plans to make money the way most open-source infrastructure companies do: a hosted version with audit logs, single sign-on and support contracts.

Leave a Reply

Your email address will not be published. Required fields are marked *

Keep reading

More stories