Hey HN! Spencer here - wanted to share a token compression tool that I’ve built for myself to save 30% costs on codex!
I've been building a ton of web apps to replace saas i'm too cheap to keep paying for like chatgpt, notion, myfitnesspal, etc. and so i can get the exact features i want (+ no ads!).
I was using Sol originally which was ok but also irritating because it would mess up and make a buggy ui.
so I switched to Astra which is awesome but it inhales my resets and leaves me with luna (no point in continuing).
I'm cheap as i said so instead of upgrading or switching to the hella expensive api (burned $700/day for similar amount of work on pro), I was thinking if I'm paying per token why not just reduce the tokens i use?
So I built a tool that shrinks the prompts and context in codex so it costs less (!) when using on codex IRL, it cut down tokens by 29.6% and now i just leave it on by default in codex and I'm actually able to get through sessions without running out of credits.
I’m making the cli free for everyone to use (https://github.com/spenmcke/compress). If you want to integrate it into your product to offer the best models at low cost, I can set you up with an sdk and api keys
PS: It’s built for coding agents, not conversational agents. I optimized it for file retrieval accuracy, trajectory preservation, and quality to get up to 30% cost reduction depending on how context-heavy the task is.
On security and privacy side, it's a proxy wrapping your local codex and zdr so it doesn't retain any of your queries. It’s on by default in codex and when you don’t want compression, you can use `codex --uncompress` to disable it.
Give it a try, would be great to hear feedback and learn your hacky ways to save token costs too!
Hey HN! Spencer here - wanted to share a token compression tool that I’ve built for myself to save 30% costs on codex!
I've been building a ton of web apps to replace saas i'm too cheap to keep paying for like chatgpt, notion, myfitnesspal, etc. and so i can get the exact features i want (+ no ads!).
I was using Sol originally which was ok but also irritating because it would mess up and make a buggy ui. so I switched to Astra which is awesome but it inhales my resets and leaves me with luna (no point in continuing).
I'm cheap as i said so instead of upgrading or switching to the hella expensive api (burned $700/day for similar amount of work on pro), I was thinking if I'm paying per token why not just reduce the tokens i use?
So I built a tool that shrinks the prompts and context in codex so it costs less (!) when using on codex IRL, it cut down tokens by 29.6% and now i just leave it on by default in codex and I'm actually able to get through sessions without running out of credits.
I’m making the cli free for everyone to use (https://github.com/spenmcke/compress). If you want to integrate it into your product to offer the best models at low cost, I can set you up with an sdk and api keys
PS: It’s built for coding agents, not conversational agents. I optimized it for file retrieval accuracy, trajectory preservation, and quality to get up to 30% cost reduction depending on how context-heavy the task is.
On security and privacy side, it's a proxy wrapping your local codex and zdr so it doesn't retain any of your queries. It’s on by default in codex and when you don’t want compression, you can use `codex --uncompress` to disable it.
Give it a try, would be great to hear feedback and learn your hacky ways to save token costs too!