4 comments

  • sarjann 2 minutes ago

    Great news, has made low memory bandwidth model usage so much nicer.

  • hypfer 24 minutes ago

    Amazing tech

    > An agent writes in an afternoon what a chatbot writes in a month

    But can you just.. not.

    Your tech is so good, it speaks for itself. Don't ruin that.

  • adefa 27 minutes ago

    I'm getting around 27 tokens per second decode using vLLM + Qwen 3.8 27b nvfp4 + DFlash 2 on the DGX Spark.

  • verdverm an hour ago