• Eager Eagle@lemmy.world
    link
    fedilink
    English
    arrow-up
    7
    ·
    22 hours ago

    Is it just me or it’s really underwhelming that they seem to remove a lot of layers and still only get 30 tokens per second? I hope they’re at least more energy efficient.

    • ☆ Yσɠƚԋσʂ ☆@lemmy.mlOP
      link
      fedilink
      arrow-up
      11
      ·
      21 hours ago

      I think it’ll depend on how much these chips cost. If it’s gonna be something like a Raspberry Pi that works as well as a 500 dollar GPU, that’s gonna be pretty amazing.

      • Asafum@lemmy.world
        link
        fedilink
        arrow-up
        7
        ·
        21 hours ago

        as well as a 500 dollar GPU

        (Cries in $1,300 Intel arc fighting tooth and nail to break 25t/s…)

        • yxc999@lemmy.ml
          link
          fedilink
          arrow-up
          2
          ·
          6 hours ago

          Huh, my B70 gets peaks of 37 tok/s TG and 800 tok/s PP, what are your settings?

          • Asafum@lemmy.world
            link
            fedilink
            arrow-up
            1
            ·
            5 hours ago

            At first I was mislead by some faulty benchmarking I was doing and thought that llama.cpp + SYCL was worse than Vulkan stock. Vulkan stock couldn’t run past 10t/s with MTP for some reason, so after a whole bunch of shenanigans I ended up again back on llama.cpp + SYCL but was lead to believe MTP was hurting me so I didn’t retry until last night. I finally got the 35-40t/s with llama.cpp + SYCL and MTP and 800PP with AOT over using JIT. I drop to 25ish after 40-50k context and stay kinda flat 12ish at 160k

            I didn’t want to move to Linux to try vLLM so this was all just trying to fight windows b.s.

            If you’re on windows and fighting with the VRAM offload after 70 seconds like I was: HKLM\SYSTEM\CurrentControlSet\Control\GraphicDrivers

            EnableRuntimePowerManagement set to 0

            • yxc999@lemmy.ml
              link
              fedilink
              arrow-up
              2
              ·
              5 hours ago

              I’m using this on Linux, so the first two lines won’t apply to you. Besides that, you are completely right, llama.cpp with the SYCL backend and MTP is the fastest way to run LLMs on Intel Arc GPUs, it’s just braindead people and bots on Reddit repeating the same line while they run a NVIDIA GPU and Ollama.

              source /opt/intel/oneapi/setvars.sh
              /opt/llama.cpp-sycl/bin/llama-server \
                --model /path/to/your/model/Qwen3.8-27B-UD-Q6_K_XL.gguf \
                --device SYCL0 \
                --n-gpu-layers 999 \
                --load-mode none \
                --flash-attn on \
                --jinja \
                --reasoning-preserve \
                --ctx-size 100000 \
                --cache-type-k q8_0 \
                --cache-type-v q8_0 \
                --temp 1.0 \
                --top-p 0.95 \
                --top-k 20 \
                --min-p 0.00 \
                --presence-penalty 0.0 \
                --repeat-penalty 1.0 \
                --spec-type draft-mtp \
                --spec-draft-n-max 2 \
                --port 9931
              
              • Asafum@lemmy.world
                link
                fedilink
                arrow-up
                1
                ·
                5 hours ago

                Thanks for the info! I believe I’m running MTP n-max 4 if I’m not mistaken which seems to be holding out well

                • yxc999@lemmy.ml
                  link
                  fedilink
                  arrow-up
                  2
                  ·
                  5 hours ago

                  I feel like you definitely can get away with as much as even 5, but if you actually look at the CLI, it will tell you how many draft tokens actually got accepted. In some situations like writing simple text, over 90% of my two draft tokens get accepted so setting it to five would have made it faster. However if you run a task which requires heavy reasoning, even two draft tokens start approaching close to only 60%. Considering Qwen3.8-27B overthinks more than an anxious teenage girl, you should generally with two. If you take my example with the quant dropped to UD-Q5_K_XL with 262K tokens and reasoning low for a task that isn’t actually complex, just long, then go right ahead and let it run with 5.

                  Edit: excuse the horrendous grammar, I was distracted