The way I have been doing it is to use LLMs to generate the code that I don't want to write: prototypes, tests, benchmarks, isolated,straightforward almost copy-paste code. I still write my own code as before because I enjoy doing that and because trying to understand and fix what an LLM generates and regenerates is harder and more tedious and time consuming than writing the code the way I want to do it in the first place.
After vibing myself into a corner multiple times on important projects, I now have only two modes: clankermaxx for code I don't really care about (mostly frontend react), and write by hand everything else. Using Django for backend already removes the most of the cruft, and writing by hand also means I actually understand what's going on. Works fine for now.
I'm still on the fence for tests: I don't really like to write them, but the LLM-generated tests are pretty bad, even the frontier models on xhigh thinking. I usually generate them, but I don't really have the confidence they test anything except 1==1. Unfortunately it's hard to justify the time spent on writing them manually.
Its a spectrum. For ai-maintained code (like a gui to visualize performance data), IDGAF what the code looks like. I just let claude or codex go nuts and 100% vibe code.
For code I care about, I audit every single hunk as its produced. I give it extensive style guidelines, and crack down on things like a 20-line essay in a comment.
For mission critical code, I write the code myself and have an agent review it.
I'm really happy with my opencode + open weight setup, the code is generally pretty good, but I do spend tokens having agents go look for common ai slop patterns.
It's heavily customized, replaced most internal systems via plugins, a set of custom agent instead of builtin ones, different model families for different sub tasks. (don't have claude review its own code)
I'm working on polishing them up and porting a few more from my own harness, then will be open sourcing. Keep your eye out for a "better-opencode" plugin suite, I'll be sure to share it with HN :]
In the near-term, I really like GLM 5.3 prose for code explore / review, give it a shot, it's cheaper (flash model) and catches all sorts of mistakes from the Big Ai models. We fully rolled out our custom pr-review on glm-5.3-flash last week. Most devs are still on claude, moving them towards fireworks and opencode.
OpenCode Go is a great way to try out open weight models for $10/month
The way I have been doing it is to use LLMs to generate the code that I don't want to write: prototypes, tests, benchmarks, isolated,straightforward almost copy-paste code. I still write my own code as before because I enjoy doing that and because trying to understand and fix what an LLM generates and regenerates is harder and more tedious and time consuming than writing the code the way I want to do it in the first place.
After vibing myself into a corner multiple times on important projects, I now have only two modes: clankermaxx for code I don't really care about (mostly frontend react), and write by hand everything else. Using Django for backend already removes the most of the cruft, and writing by hand also means I actually understand what's going on. Works fine for now.
I'm still on the fence for tests: I don't really like to write them, but the LLM-generated tests are pretty bad, even the frontier models on xhigh thinking. I usually generate them, but I don't really have the confidence they test anything except 1==1. Unfortunately it's hard to justify the time spent on writing them manually.
Its a spectrum. For ai-maintained code (like a gui to visualize performance data), IDGAF what the code looks like. I just let claude or codex go nuts and 100% vibe code.
For code I care about, I audit every single hunk as its produced. I give it extensive style guidelines, and crack down on things like a 20-line essay in a comment.
For mission critical code, I write the code myself and have an agent review it.
I'm really happy with my opencode + open weight setup, the code is generally pretty good, but I do spend tokens having agents go look for common ai slop patterns.
It's heavily customized, replaced most internal systems via plugins, a set of custom agent instead of builtin ones, different model families for different sub tasks. (don't have claude review its own code)
I'm working on polishing them up and porting a few more from my own harness, then will be open sourcing. Keep your eye out for a "better-opencode" plugin suite, I'll be sure to share it with HN :]
In the near-term, I really like GLM 5.3 prose for code explore / review, give it a shot, it's cheaper (flash model) and catches all sorts of mistakes from the Big Ai models. We fully rolled out our custom pr-review on glm-5.3-flash last week. Most devs are still on claude, moving them towards fireworks and opencode.
OpenCode Go is a great way to try out open weight models for $10/month