llama.cpp b10975: cuda : enable i16 and i32 for DUP (#28897)

cuda : enable i16 and i32 for DUP docs : update ops table for DUP on CUDA