ArXivCSExplorer
☆☆Bookmarks🏆RSSHow to UseFAQ
Built with and by Teycir Ben Soltane•
How to Use•FAQ•GitHub•arXiv.org•
Share:

20 results for “Understanding of serverless computing, large language models, energy optimization, scheduling”

CS papers only

Hybrid search: Keyword + semantic, ranked by combined score.ⓘ

Want pure semantic search? Try claim verification →

cs.LGcs.AIcs.DCEmpiricalRecentJul 16, 2026

An Auto-Scaling Approach for Serverless Environments Based on a Multi-Expert Consensus Mechanism

Mobina Kashaniyan, Mehrdad Ashtiani, Amirhossein Ghassemi

This paper proposes a dependency-aware autoscaling framework for serverless computing, integrating graph-based bottleneck identification, short-term workload forecasting, multi-model consensus, and co…

View →
cs.DCEmpiricalRecentJun 29, 2026

Energy-Aware Scheduling for Serverless LLM Serving on Shared GPUs

Tianyu Wang, Gourav Rattihalli, Aditya Dhakal, Longfei Shangguan +1 more

This paper presents Festina, a profiling-guided, power-aware control plane for minimizing energy consumption in serverless large language model (LLM) serving.

View →
cs.CRcs.AIcs.CLRecentMar 17, 2026

Resource Consumption Threats in Large Language Models

Yuanhe Zhang, Xinyue Wang, Zhican Chen, Weiliu Wang +7 more

This survey systematically reviews resource consumption threats in large language models (LLMs) to provide a unified view of the problem landscape, from threat induction to mitigation.

View →
cs.CRcs.DCRecentMar 20, 2026

Kumo: A Security-Focused Serverless Cloud Simulator

Wei Shao, Khaled Khasawneh, Setareh Rafatirad, Houman Homayoun +1 more

The paper introduces Kumo, a novel security-focused simulator that enables controlled analysis of resource sharing and scheduling risks in serverless cloud environments, demonstrating that scheduler c…

View →
cs.OScs.ARcs.NIEmpiricalRecentJul 17, 2026

Rethinking Polling Efficiency in Service Core Network Stacks

Matheus Stolet, Simon Peter, Antoine Kaufmann

This paper argues that idle cores on contemporary multicore processors can return compute capacity and proposes a budget-centric view of service core systems.

View →
cs.CRRecentMar 26, 2026

ALPS: Automated Least-Privilege Enforcement for Securing Serverless Functions

Changhee Shin, Bom Kim, Seungsoo Lee

ALPS is an automated, vendor-agnostic framework that enforces least privilege in serverless functions by analyzing code and generating precise security policies, achieving high coverage and significan…

View →
cs.DMTheoreticalRecentJul 3, 2026

Scheduling Tasks towards Energy Autarky: Benefits and Computational Costs of Flexibility

Robert Bredereck, Till Fluschnik, Klaus Heeger

This paper studies the autarky problem of scheduling energy-consuming jobs with time windows using a battery and an energy forecast, and shows NP-hardness, polynomial-time solvability, and fixed-param…

View →
cs.LGcs.DCEmpiricalRecentJul 24, 2026

SLA-Constrained Carbon-Aware Routing in Geo-Distributed Serverless Clouds

Anmol Chaudhary, Rahul Mishra

This paper proposes a carbon-aware routing policy for geo-distributed cloud deployments, achieving up to 46.8% carbon reduction while maintaining zero SLA violations.

View →
cs.DCEmpiricalRecentJun 29, 2026

Spandana: Reconciling Strict SLOs with Low Cost under Fine-Grained Load Fluctuations

Dilina Dehigama, Shyam Jesalpura, Zeyu Xu, Marton Nemeth +3 more

The paper introduces Spandana, an architecture that decouples SLO enforcement from cost optimization in cloud-based online services, achieving high utilization, strict SLO adherence, and cost savings.

View →
cs.ETcs.DCEmpiricalRecentJul 21, 2026

Examining QRMI as a Unified Interface for Quantum-HPC Integration

Thomas Badts, Tim Boyle, Claudio Carvalho, Antonio Córcoles +24 more

The paper presents the Quantum Resource Management Interface (QRMI) as a standardized, vendor-agnostic middleware layer for integrating quantum resources into high-performance computing environments,…

View →
cs.AIcs.DBRecentMay 27, 2026

A Query Engine for the Agents

Kenny Daniel

The paper introduces Hyperparam, a set of lightweight JavaScript libraries designed to enable direct, model-aware querying of unstructured data (like agent traces) within client-side AI applications.

View →
cs.ARcs.LGEmpiricalRecentJun 11, 2026

BigPower: Hierarchical Source-Level Module Power Estimation for CPUs with Large Language Models

Honghua Zhu, Chunjie Luo, Jianfeng Zhan

This paper introduces BigPower, a hierarchical source-level surrogate model for fine-grained module-level power estimation during CPU design using large language models and architectural hierarchy.

View →
cs.DCcs.AIcs.OSEmpiricalRecentJul 2, 2026

Fine-Grained Computation Offload for Off-the-Shelf Servers in Tens of Lines

Bojie Li

The paper proposes a method to improve the performance of fine-grained offloads on servers by overlapping the offload with other requests using server-side routing.

View →
cs.PLcs.LGTheoreticalRecentJul 23, 2026

Relaxed activation analysis of dataflow networks - A clock calculus for machine learning and real-time scheduling

William Gaudelier, Albert Cohen, Dumitru Potop Butucaru

The paper proposes a conservative extension of Lustre's clock calculus to facilitate the embedding of ML models in reactive applications.

View →
cs.CLcs.AIRecentMay 29, 2026

D$^3$: Dynamic Directional Graph-Constrained Data Scheduling for LLM Training

Yuanjian Xu, Jianing Hao, Guang Zhang, Zhong Li

The paper proposes $D^3$, a dynamic graph-constrained scheduling framework that optimizes LLM training order by modeling sample interactions as a dynamic influence graph.

View →
cs.SEcs.CLSurveyRecentJun 18, 2026

Token-Operations-Oriented Inference Optimization Techniques for Large Models

Shiguo Lian, Kai Wang, Zhaoxiang Liu, Wen Liu +21 more

This paper proposes a four-layer technical architecture for large model inference optimization, including Multi-model Fusion, Model Optimization, Compute-Model Fusion, and Compute-Network-Model Fusion…

View →
cs.DCEmpiricalRecentJun 29, 2026

Towards Transparent Checkpointing with AI-driven Code Generation

Hai Duc Nguyen, Tekin Bicer, Kyle Chard, Ian Foster +1 more

Researchers used a large language model to generate checkpoint/restart code for MPI applications, achieving comparable efficiency to human-engineered solutions.

View →