Look at the PDF in simonw’s first link. There’s plenty of information. One part looked like the communication requirements were reduced down to under 100MB. That suggests a communication rate that could be handled by dirt-cheap instances spread across the globe. Like on vast.ai or something.
GaLore can do that too (if you transplant it). There are similar methods prior LLM era doing the same. They are not quite there on the loss graph side though (I am actually unsure about GaLore, but both FLoRA and ReLoRA were not quite there on loss graph side).