Skip to content

Repository files navigation

MiniSky

Lightweight cloud orchestration tool inspired by SkyPilot. Run your machine learning and data science workloads easily on multiple cloud providers (RunPod, Lambda Cloud, AWS, GCP) with a single command.

Features

  • Multi-Cloud Support: Deploy to RunPod, Lambda Cloud, AWS EC2, GCP Compute Engine, or use the Mock provider for testing.
  • Cost Optimizer: Automatically selects the cheapest provider for your requirements.
  • Autostop: Every launched VM is watched by a detached background process and stopped after 30 idle minutes by default (override per-task with autostop_minutes, or skip it entirely with --no-autostop) - protects against runaway bills even after minisky launch returns or the terminal is closed.
  • File Synchronization: Automatically sync local directories (workdir) to remote VMs.
  • Managed Jobs: minisky jobs launch automatically relaunches a task if its (spot) instance is preempted, with checkpoint save/restore.
  • Web Dashboard: minisky serve runs a FastAPI + Vue.js dashboard for creating/monitoring clusters through a browser instead of the terminal. VMs launched via minisky launch/minisky cluster launch show up there too, and vice versa - see Web Dashboard below.

Installation

MiniSky isn't published to PyPI yet — install from source with uv:

git clone https://github.com/berketpbs/MiniSky.git
cd MiniSky
uv venv
uv pip sync pyproject.toml

Quick Start

  1. Set up your cloud credentials in ~/.minisky/config.yaml:
providers:
  runpod:
    api_key: "YOUR_RUNPOD_KEY"
  lambda:
    api_key: "YOUR_LAMBDA_KEY"
  aws:
    # Optional - if omitted, falls back to the standard AWS credential
    # chain (`aws configure`, AWS_ACCESS_KEY_ID/AWS_SECRET_ACCESS_KEY, IAM role)
    access_key_id: "YOUR_AWS_ACCESS_KEY_ID"
    secret_access_key: "YOUR_AWS_SECRET_ACCESS_KEY"
    region: "us-east-1"
    key_name: "your-ec2-keypair-name"       # required to SSH into launched instances
    security_group_id: "sg-xxxxxxxx"        # optional - must allow inbound SSH (22)
  gcp:
    project: "your-gcp-project-id"          # required - GCP has no default project
    # Optional - if omitted, falls back to google-auth's standard chain
    # (GOOGLE_APPLICATION_CREDENTIALS, `gcloud auth application-default login`,
    # or the GCE metadata server)
    credentials_path: "/path/to/service-account.json"
    zone: "us-central1-a"
    ssh_public_key_path: "~/.ssh/id_rsa.pub"  # required to SSH into launched instances
  1. Create a task YAML file (task.yaml):
name: my-training-job
resources:
  gpu: "A100"
  gpu_count: 1
  disk_gb: 100
workdir: ./src
run:
  - python train.py
  1. Launch your task:
minisky launch task.yaml

Note: the mock provider returns 127.0.0.1 as the VM's IP but doesn't run a fake SSH server — MiniSky still opens a real SSH connection to execute setup/run. To see a task actually execute end-to-end with mock, you need a local SSH server listening on port 22 (e.g. Windows OpenSSH Server, or any Linux/macOS machine has one by default). Without one, launch will time out at the "Waiting for SSH" step — state tracking, the CLI, and providers can still be exercised, just not remote execution.

CLI Reference

VM lifecycle

minisky launch task.yaml          # Launch a VM and run the task
minisky launch task.yaml --no-autostop   # Skip autostop for this VM
minisky status [vm-id]            # Show status of one or all VMs
minisky stop <vm-id>              # Stop a VM, preserving disk
minisky start <vm-id>             # Start a previously stopped VM
minisky terminate <vm-id>         # Terminate a VM and clean up state

Working with a running VM

minisky exec <vm-id> "nvidia-smi"           # Run a command remotely
minisky ssh <vm-id>                         # Open an interactive SSH session
minisky port-forward <vm-id> jupyter tensorboard   # Forward ports locally
minisky logs <vm-id> -f                     # Stream logs
minisky sync <vm-id> ./local --remote ~/remote     # Sync files to/from a VM (rsync/SFTP)
minisky rsync <vm-id> ./local ~/remote             # Quick rsync shortcut

Fleet-level

minisky check                     # Verify setup and provider credentials
minisky gpus                      # Browse GPU pricing/availability across providers
minisky optimize task.yaml         # Rank providers/prices for a task without launching
minisky cost-report                # Cost report across all tracked VMs
minisky config show|set|get|unset  # Manage ~/.minisky/config.yaml
minisky cluster launch|status|terminate   # Multi-node cluster management
minisky queue list|add|show|cancel|clear  # Job queue management
minisky jobs launch|list|status|cancel    # Managed jobs: auto-relaunch on spot preemption

Run minisky <command> --help for full options on any command.

Web Dashboard

minisky serve                 # Starts the API on :8000 (and the UI too, if built - see below)

By default this serves only the FastAPI backend (REST + WebSocket at /v1). To also serve the Vue web UI from that same process:

cd dashboard
npm install
npm run build
cd ..
minisky serve                 # now also serves the built UI at http://localhost:8000/

For frontend development with hot-reload, run the dashboard's own dev server instead (cd dashboard && npm run dev, on :3000) alongside minisky serve.

To bind the API to a different interface or port, pass --host and --port:

minisky serve --host 127.0.0.1 --port 8080

Note: clusters are shared between the CLI and the dashboard - a VM launched via minisky launch/minisky cluster launch shows up in the dashboard (tagged CLI there) and can be stopped/started/terminated from it, and a cluster created via the dashboard shows up in minisky status too. Jobs are not shared yet: the dashboard's job submission/execution (minisky/api/core.py's JobController) is a separate system from minisky queue/minisky jobs on the CLI side.

Development

Use uv to install dependencies and run tests:

uv venv
uv pip sync pyproject.toml
uv run pytest tests/

Run a focused test module while developing:

uv run pytest tests/test_config.py

About

Seamless cloud orchestration for AI/ML. Deploy and manage GPU workloads across multiple clouds effortlessly.

Resources

Contributing

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages