Formalization and quantitative metrics for functional stability of edge computing systems
This paper proposes a framework for functional stability in edge computing systems, focusing on individual functions and their quality thresholds.
The paper proposes a new approach to reliability and fault-tolerance in edge computing systems, shifting the focus from the system to individual functions.
Before reading this…
Applications
- →UAV swarms, edge AI inference services
To understand this paper, make sure you know these concepts first:
- Understanding of edge computing systems, basic knowledge of reliability and fault-tolerance conceptsfind papers →
Abstract
More Like ThisEdge computing systems operate in conditions where component failures, intermittent backhaul connectivity, resource exhaustion and adversarial disturbances are part of the normal operating envelope rather than rare events. Classical reliability and fault-tolerance models, oriented toward monolithic systems and binary working/failed states, do not capture the differentiated quality requirements of services co-located on a resource-constrained edge node. This article proposes a formal framework of functional stability that shifts the unit of analysis from the system to the individual function. Strong and weak forms of stability are defined through a continuous quality function and per-function quality thresholds. A system of base parameters (disturbance tolerance, recovery time, degradation depth, admissible degradation) and aggregate metrics (cumulative quality, functional availability, integral stability) is introduced and connected to the formal definitions by nine theorems. The framework is instantiated for edge computing scenarios and illustrated through two analytical case studies: a UAV swarm and an edge AI inference service with cloud fallback. The proposed metrics distinguish architectural alternatives that classical availability cannot separate and expose the dependence of architectural ranking on the explicitly selected disturbance-class set.