3 comments

  • Marcuss2 41 minutes ago

    There is something wrong about this.

    > The Planner, Reviewer and Tester agents were assigned the same configuration in every run: claude-opus-4.8 via claude-code, only the Builder configuration was dynamic in the team setup. So, the results read as the difference between builders. That's a SWE in a team.

    Yet they show the total cost. Making the comparison irrelevant. Most of the cost is going to be in there with model that expensive.

    Also, they show open weight models in `opencode` (I have reservations about it, but fine) and `pi`, in it's most basic config (who actually uses it that way with models this capable?).

    Show me actual comparison, from my experience, the planner, reviewer and tester should be different models from different families.

      onder_ceylan 27 minutes ago

      I understand where you're coming from. The research focuses on Builder configuration being variable and whether it affects metrics like cycle time and cost, which it shows the diff in detail.

      I'm interested to see the total inference cost of a ticket type, that's why total cost is compared. Even though most of the cost is there in other stages, there's still huge difference even with one variable configuration in the loop.

      The cost of each builder is also shown in mission breakdown, but good feedback for extending some of the tables, thanks!

        Marcuss2 3 minutes ago

        Replace all of the engineers with open weight models. And add GLM 5.3 Flash as a builder. I am interested how will they fare.