Author here. We tested a simple controller designed to reduce pauses while an LLM generates a response.
In one setup, it reduced P99 inter-token latency by 27.7%. But it did not work on larger models or multiple GPUs. We traced the failure to the timing signal used by the controller.
We published both the positive and negative results because the failure identifies an important limitation and suggests what a better controller should measure.
Happy to answer questions about the implementation, experiments, or results.
Author here. We tested a simple controller designed to reduce pauses while an LLM generates a response. In one setup, it reduced P99 inter-token latency by 27.7%. But it did not work on larger models or multiple GPUs. We traced the failure to the timing signal used by the controller. We published both the positive and negative results because the failure identifies an important limitation and suggests what a better controller should measure.
Happy to answer questions about the implementation, experiments, or results.