llama.cpp b10988

opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP ( #27637 ) opencl: gate the prebuilt q4_0 MoE GEMM on routing count…

opencl: choose the MoE expert matmul by batch size for speculative decoding/MTP ( #27637 ) opencl: gate the prebuilt q4_0 MoE GEMM on routing count opencl: stop writing zeros into the padded MoE activation slots opencl: rephrase claude's comments Co-authored-by: Li He…

Read the original source — github.com

release · Shared by tscosj

0 comments

No comments yet.