Mouse scored 1st on the Frontier Harness benchmark last night and Verification loops are a big reason why.
The benchmark runs Kimi K3 through leading coding harnesses across a range of software tasks to measure how much the harness itself affects performance.
My main takeaway from our results is that verification loops and a few deterministic rules can make a coding harness perform significantly better, even with the exact same model.
Mouse scored 1st on the Frontier Harness benchmark last night and Verification loops are a big reason why.
The benchmark runs Kimi K3 through leading coding harnesses across a range of software tasks to measure how much the harness itself affects performance.
My main takeaway from our results is that verification loops and a few deterministic rules can make a coding harness perform significantly better, even with the exact same model.
More in the blog.