Z.ai says its week-long anonymous GLM-5.3-Flash preview ran entirely on domestic Chinese accelerators, at per-token cost it calls comparable to Nvidia GPUs. It named no chip vendor and released no throughput or power figures. Serving is also the easier half of the problem.
Wait can tokens just be directly compared like that? My impression is that token cost can vary by 2 orders of magnitude, depending on model, because the actual work of computation varies by that much between models.