llama.cpp b10981

OpenVINO: optimize stateful decode and GPU MoE inference ( #28638 ) exclude GPU/NPU failing POOL_2D case Fix pool case ggml-openvino: fix stateful decode…

OpenVINO: optimize stateful decode and GPU MoE inference ( #28638 ) exclude GPU/NPU failing POOL_2D case Fix pool case ggml-openvino: fix stateful decode for Gemma-4 per-layer-type head sizes ggml-openvino: fix MSVC narrowing error in permute ggml-openvino: classify…

Read the original source — github.com

release · Shared by tscosj

0 comments

No comments yet.