20 results for “latency variations”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
This paper investigates a network side-channel vulnerability in multi-tenant datacenter fabrics caused by shared congestion behavior, achieving up to 97.3% run-level accuracy in workload inference.
This paper proposes a hierarchical analytical framework to characterize region-level latency differences in Low-Earth orbit satellite Internet using Starlink RTT measurements.
This paper proposes a new approach for designing Radio Access Network (RAN) slices in 5G and beyond networks using descriptors that consider both transmission rate and latency requirements to support…
This paper proposes CAPS, a scheduling layer for data centers that separates rate computation and packet scheduling, reducing queue occupancy by up to 10x without throughput loss.
Siyuan Shen, Anton Korzh, John Bachan, Tiancheng Chen +9 more
This paper explores methods to reduce latency in GPU collective communications for large language model inference, achieving near-optimal designs with barrier-free synchronization and efficient use of…
The paper proposes a multi-resolution end-to-end deep neural network for autonomous driving that dynamically adjusts input resolution to optimize the critical tradeoff between prediction accuracy and…
Mohammadparsa Karimi, Majid Nabi, Ahmed Khalaf, Andrew Nelson +2 more
This paper introduces Virtual Time-Sensitive Networking (V-TSN), a software-defined overlay for gPTP-based synchronization and TSN traffic shaping over general-purpose networks without specialized har…
This paper investigates the latency performance of Mobile Edge Computing (MEC) on a 5G cellular network for real-time power transmission line analytics, demonstrating a low latency of 44.62 ms compare…
The paper implements and evaluates several hardware priority queue architectures on modern FPGA platforms and provides a quantitative analysis.
This paper evaluates user-perceived performance and QoE of wireless networks at Notre Dame Stadium during football games using commercial smartphones for web browsing, WhatsApp messaging, and Instagra…
Gatling is a new atomic broadcast protocol that achieves arbitrarily small inter-proposal times, even smaller than the network delay, by running multiple parallel instances of a black-box atomic broad…
This paper compares Google's BBR-v3 Congestion Control Algorithm to eight others over SpaceX's Starlink network, demonstrating its fairness and throughput maximization in high-latency, variable satell…
This paper analyzes the effects of calibration errors on downlink beamforming in a repeater-assisted massive MIMO system and derives analytical expressions for the downlink spectral efficiency.
This paper proposes a proactive fair scheduling framework to prevent fabric oversubscription and eliminate incast in Mixture of Experts (MoE) architectures, demonstrating consistent link utilization a…
Kalle Kujanpää, Ning Liu, Shahnawaz Alam, Yeshwanth Reddy Sura +3 more
The paper presents a tool-making pipeline for production LLM agents that compiles repeated steps into validated, versioned tools before deployment, reducing latency and error rate.
Guanjie Lin, Yinxin Wan, Shichao Pei, Ting Xu +2 more
The paper introduces GateScope, a black-box framework that audits commercial LLM API gateways, revealing frequent discrepancies in model behavior, billing, and performance across real-world services.