Score: 0

SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost

Published: January 10, 2026 | arXiv ID: 2601.06520v1

By: Zhifei Li , Tian Xia , Ziming Mao and more

Potential Business Impact:

Saves money on computer jobs by using cheaper, temporary power.

Business Areas:

Cloud Computing Internet Services, Software

AI batch jobs such as model training, inference pipelines, and data analytics require substantial GPU resources and often need to finish before a deadline. Spot instances offer 3-10x lower cost than on-demand instances, but their unpredictable availability makes meeting deadlines difficult. Existing systems either rely solely on spot instances and risk deadline violations, or operate in simplified single-region settings. These approaches overlook substantial spatial and temporal heterogeneity in spot availability, lifetimes, and prices. We show that exploiting such heterogeneity to access more spot capacity is the key to reduce the job execution cost. We present SkyNomad, a multi-region scheduling system that maximizes spot usage and minimizes cost while guaranteeing deadlines. SkyNomad uses lightweight probing to estimate availability, predicts spot lifetimes, accounts for migration cost, and unifies regional characteristics and deadline pressure into a monetary cost model that guides scheduling decisions. Our evaluation shows that SkyNomad achieves 1.25-3.96x cost savings in real cloud deployments and performs within 10% cost differences of an optimal policy in simulation, while consistently meeting deadlines.

Deadline-Aware Online Scheduling for LLM Fine-Tuning with Spot Market Predictions

Distributed, Parallel, and Cluster Computing

Saves money training big computer brains.

24 Dec 2025 0

86%

Simulating Dynamic Cloud Marketspaces: Modeling Spot Instance Behavior and Scheduling with CloudSim Plus

Distributed, Parallel, and Cluster Computing

Saves money on computer clouds by handling price changes.

22 Nov 2025 1

86%

Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications

Distributed, Parallel, and Cluster Computing

Trains AI faster on different computer parts.

24 Dec 2025 0

View PDF Login to Bookmark

Page Count

17 pages

SkyNomad: On Using Multi-Region Spot Instances to Minimize AI Batch Job Cost

Saves money on computer jobs by using cheaper, temporary power.

Technical Abstract

Deadline-Aware Online Scheduling for LLM Fine-Tuning with Spot Market Predictions

Simulating Dynamic Cloud Marketspaces: Modeling Spot Instance Behavior and Scheduling with CloudSim Plus

Diving into 3D Parallelism with Heterogeneous Spot Instance GPUs: Design and Implications