ProFlow: RL-Driven and Performance-Aware Proactive Flow Placement in Datacenter Networks
Proposed ProFlow, a proactive flow-placement framework using distributed telemetry signals and offline-trained RL for identifying precursor congestion conditions and rerouting protected flows in multi-tenant datacenter networks, achieving approximately 40% higher mean throughput and initiating rerouting decisions around 34 seconds earlier on average.
Proposes a proactive flow-placement framework using distributed telemetry signals and RL for early congestion management
Before reading this…
Applications
- →Multi-tenant datacenter networks
To understand this paper, make sure you know these concepts first:
- Understanding of datacenter networks and congestion managementfind papers →
Abstract
More Like ThisIn datacenter fabrics composed of leaf and aggregation switches, competing flows may become co-located on shared aggregation switches, creating congestion that can significantly degrade protected flows. However, before throughput degradation becomes observable, the network often exhibits early signs characterized by rising flow activity and queue overflow signals. Existing congestion-management approaches primarily react only after congestion becomes visible, leaving these early signs largely unexploited. In this paper, we propose ProFlow, a proactive flow-placement framework for protecting performance-sensitive traffic in multi-tenant datacenter networks, thereby utilizing the early signs of potential throughput degradations. ProFlow leverages distributed telemetry signals and offline-trained reinforcement learning (RL) to identify precursor congestion conditions and proactively reroute protected flows before throughput degradation occurs. Evaluation results using FABRIC testbed show that ProFlow achieves approximately 40% higher mean throughput than a reactive rerouting baseline while initiating rerouting decisions around 34 seconds earlier on average, demonstrating the effectiveness of anticipatory congestion management.