TP: fix split state and granularity for fused QKV gemma4, qwen35 ( #28965 ) model: calculate split states for attn_qkv from n_head * n_embd_head_k required for gemma4 with --fuse-qkv, where n_embd is 5376 but Q is 8192. model: handle fused full attention layers for…
Read the original source — github.com
release · Shared by tscosj
0 comments
No comments yet.