12 comments

  • msdz 4 minutes ago

    Congratulations on the launch, it looks like an impressive product and tool!

    Q: From my (very, very limited!) understanding, I was under the impression that part of the “inference engine inertia” is model- or at least architecture-specific code for most, if not each new open-weight model coming out. Assuming I got that right, do you support, or plan on supporting, everything vLLM/llama.cpp do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?

  • lxe 7 minutes ago

    On my local inference box I have a perpetual codex thread open in my llama.cpp checkout that I periodically ask to take a look at currently pending llama.cpp PRs, do some research on latest MTP, Dflash and other prediction or attention optimizations, do research on the latest model quants and finetunes, take a look at localLlama Reddit threads and just do essentially a sweep of the frontier.

    Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.

    Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.

  • kmike84 12 minutes ago

    This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)

    I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.

    3 main failure modes I observed in the engines:

    * Not using best available spec decoding

    * Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)

    * Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline

  • nateb2022 28 minutes ago

    Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.

      anerli a minute ago

      The benchmark we cited here is a simple prose-repetition task. We put the content of Moby Dick up to 64k context in the request, and then ask it to repeat the last section.

      For llama.cpp, we try to make the comparison as fair as possible by using similar settings. No speculative decoding, default prefill batch sizes, flash attention on.

      We tried also quantizing the KV cache to 8-bit keys and 4-bit values like we do in Magnitude, but this bombed decode speed for llama.cpp in our testing. Since it seems llama.cpp did not optimize that path, we used 16-bit KV instead.

      The source for the benchmark is available here also: https://github.com/magnitudedev/magnitude/tree/main/inferenc...

  • sgtwompwomp 27 minutes ago

    This is dope, is this kind of like Wafer.ai but for local models? As in a coding agent optimizes the kernels so the local model runs continuously better? Cause that is compelling if so. If it’s more simple that’s cool too

  • kenzic 25 minutes ago

    How long does tuning take (on an M3 MacBook Pro for example)?

      anerli 9 minutes ago

      Tuning is a one-time process that takes around ~1 minute whenever you download a new model. This is generally enough time to tune all the kernels' parameters to the point where tuning any longer asymptotes. Time can vary a little based on the hardware though.

  • amirhesham 29 minutes ago

    Oh this is so cool. Curious about the business model, too.

  • p-e-w 35 minutes ago

    What is the business model?

      anerli 13 minutes ago

      We envision a future where workloads are hybrid. Average consumer hardware will be able to handle a lot with local models, but you’ll still want to use cloud models for harder tasks. Magnitude will make it seamless to switch between the two, even for the same tasks (without breaking your prefix cache). We’ll charge per token for our inference cloud, using the same efficiencies we unlock for local inference to pass the savings on to you.

  • yolandac 21 minutes ago

    does it allow us to run larger models that weren't possible before?