20 results for “stochastic load balancing”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
A new stochastic load balancing model is introduced that allows for a tradeoff between non-adaptive policies and performance, with results showing a 2-reservation approximation to the omniscient optim…
Dilina Dehigama, Shyam Jesalpura, Zeyu Xu, Marton Nemeth +3 more
The paper introduces Spandana, an architecture that decouples SLO enforcement from cost optimization in cloud-based online services, achieving high utilization, strict SLO adherence, and cost savings.
This paper proposes a dependency-aware autoscaling framework for serverless computing, integrating graph-based bottleneck identification, short-term workload forecasting, multi-model consensus, and co…
This paper argues that idle cores on contemporary multicore processors can return compute capacity and proposes a budget-centric view of service core systems.
This paper proposes a splay-like rotation design for concurrent binary search trees to preserve the main benefit of splaying on skewed workloads while reducing contention near the root.
This paper proposes a proactive fair scheduling framework to prevent fabric oversubscription and eliminate incast in Mixture of Experts (MoE) architectures, demonstrating consistent link utilization a…
The ZCube topology, which eliminates path multiplicity and reduces switching hardware, delivers better performance for large model training and inference than traditional multipath datacenter networks…
This paper proposes a host-driven method for flowlet balancing using Segment Routing over IPv6 (SRv6), reducing tail latency by 15% and 33% compared to random flowlet balancing and ECMP, respectively.
The paper extends results for interval scheduling to the more general throughput problem in the real-time model with constant competitive ratios for specific weight functions and advance notice.
The I/O Resource Manager (IORM) is presented as a multi-stage distributed scheduler to maintain consistent performance and enforce global I/O limits in shared-nothing disaggregated storage clusters.
Zhiyao Xu, Aoxue Liu, Zhanjie Ding, Dan Zhao +2 more
The paper proposes Task-Aware Coactivation Grouping (TACG) to significantly reduce communication costs in multi-task MoE inference by grouping experts based on task-specific co-activation patterns, ou…
This paper proposes LMS-AR, a memory bandwidth regulation mechanism for multi-core systems using a Linux kernel module with adaptive filtering for prediction and regulation.
This paper introduces Stigmergic Graph Memory (SGM), a method to improve warehouse throughput in many-to-many Multi-Agent Pickup and Delivery (MAPD) by using a bounded, decaying memory layer to record…