Author here. I wrote this because I wanted my private inference boxes available from other devices. Today I dogfooded the guide by serving the new Qwen3.8-27B Q4 on an old 8GB Quadro P4000 + 128GB RAM box.
It uses an OpenAI-compatible API and can power private apps or agentic sessions; I’m currently using local models for creative writing and business projects. Qwen3.8 still needs some per-machine tuning and can use a lot of thinking tokens, but in one particularly extreme test it one-shot an absolutely gorgeous app after about 19 hours of thinking.
Author here. I wrote this because I wanted my private inference boxes available from other devices. Today I dogfooded the guide by serving the new Qwen3.8-27B Q4 on an old 8GB Quadro P4000 + 128GB RAM box.
It uses an OpenAI-compatible API and can power private apps or agentic sessions; I’m currently using local models for creative writing and business projects. Qwen3.8 still needs some per-machine tuning and can use a lot of thinking tokens, but in one particularly extreme test it one-shot an absolutely gorgeous app after about 19 hours of thinking.