Z.ai says its week-long anonymous GLM-5.3-Flash preview ran entirely on domestic Chinese accelerators, at per-token cost it calls comparable to Nvidia GPUs. It named no chip vendor and released no throughput or power figures. Serving is also the easier half of the problem.
Yeah this is a bad metric. GLM outputs a lot of reasoning tokens.