Every Microsecond Matters: Achieving Near Speed-of-Light Latency in GPU Collectives | ArxivCSExplorer