20 results for “datacenter networks, ZCube topology, Braess's paradox, large model training, inference”
CS papers onlyHybrid search: Keyword + semantic, ranked by combined score.ⓘ
Want pure semantic search? Try claim verification →
The ZCube topology, which eliminates path multiplicity and reduces switching hardware, delivers better performance for large model training and inference than traditional multipath datacenter networks…
This paper critically re-evaluates the use of Graph Neural Networks (GNNs) for Bitcoin fraud detection, demonstrating that under strict, leakage-free temporal evaluation, simple feature-only models si…
Taurus is a single-machine system for efficient Graph Neural Network (GNN) inference on large-scale graphs that do not fit in RAM, using source-centric broadcasts and a pipelined GPU-CPU-SSD hierarchy…
This paper introduces approximation-preserving coresets, which provide weaker guarantees than strong coresets but stronger guarantees than weak coresets for preserving the costs of good solutions in b…
This paper develops a runtime recovery framework for broadcasting in dense Gaussian networks under static and dynamic faults, proving necessary and sufficient repair edges and providing efficient repa…
This paper evaluates the feasibility and cost-effectiveness of large-scale AI data centers in low-Earth orbit (LEO) versus terrestrial facilities, considering factors like launch cost, power generatio…
The paper introduces and explores Truly Linear FPT (TLFPT), a complexity class defined by $O(n) + f(k)$, demonstrating that it is a strict subset of standard Linear FPT and providing new algorithms fo…
This paper proposes an HPC-aware methodology for Knowledge Distillation (KD) that decouples teacher and student partitioning efficiently, achieving up to 67% higher samples-per-second than the widely…
This paper proposes a budget-adaptive routing method for edge-cloud inference collaborations, which selects between weak-skipping and weak-conditioned placement based on offload budget.
This paper proposes and evaluates the integration of Federated Learning and blockchain technology over cloud-edge infrastructure to enhance data privacy and security for decentralized AI applications.
This paper studies adversarial attacks on programming-by-example systems and introduces a defense method called version-space partition aggregation (VPA).
The paper extends modular dynamic Bayesian networks (MDBNs) to model non-Markovian queues, providing the first causal metamodeling technique for such systems with significant speedup.
Shiguo Lian, Kai Wang, Zhaoxiang Liu, Wen Liu +21 more
This paper proposes a four-layer technical architecture for large model inference optimization, including Multi-model Fusion, Model Optimization, Compute-Model Fusion, and Compute-Network-Model Fusion…
The paper reframes Parameter-Efficient Fine-Tuning (PEFT) from a mere cost-saving alternative to a robust architecture for creating persistent, personalized models that layer specific behaviors onto l…
The paper introduces Rotary GPU, an exploratory execution approach demonstrating that large Mixture-of-Experts models can be run locally on consumer GPUs with limited VRAM, achieving usable decode thr…
Stephan Krenn, Omid Mir, Thomas Lorünser, Sebastian Ramacher +1 more
The paper proposes a provably secure path validation protocol for large-scale Quantum Key Distribution (QKD) networks that allows receivers to verify network compliance without revealing sensitive top…
This paper studies a random compositional model for the growth of affine regions in deep piecewise-linear networks and proves the existence of a submultiplicative pressure for the number of affine pie…