• Kestrel Labs’ open-weights model just beat the incumbents at code review. The incumbents are not amused.

    On Tuesday morning, a 40-person company in Oslo did something its far larger rivals have spent a year saying was too risky. Kestrel Labs published the full weights of K2-Review, a model trained only to read code changes and comment on them, under the Apache 2.0 license.

    By the afternoon, three independent teams had run it against their own review queues. All three told Modelodge it caught more real bugs than the closed models they pay for today.

    What K2-Review actually does

    K2-Review is small by current standards: 34 billion parameters, fine-tuned on 11 million annotated pull requests. It does not write code. It reads a diff, the surrounding files and the ticket, then leaves line-level comments the way a senior engineer would.

    The benchmark that matters

    Kestrel’s headline number comes from a test set of 4,200 merged pull requests that were later reverted because they caused an incident. K2-Review flagged 71% of them. The best closed model in the comparison flagged 58%.

    Nobody gets promoted for the bug that didn’t ship. That’s why review never got the attention generation did.

    Ingrid Solberg · CTO, Kestrel Labs

    What happens next

    Kestrel plans to make money the way most open-source infrastructure companies do: a hosted version with audit logs, single sign-on and support contracts.