Tối ưu phục vụ mô hình lớn Kimi K-series và GLM trên Workers AI bằng lượng tử hóa và kiểm tra toàn vẹn KV cache — Cloudflare Insights