> ## Documentation Index
> Fetch the complete documentation index at: https://docs.basaltic.sh/llms.txt
> Use this file to discover all available pages before exploring further.

# Quotas and troubleshooting

> Why a launch failed, what the two state fields are each telling you, the three counters an instance holds at once, and the refusals worth recognising.

## Desired state and current state

Read `current_state` to see where the instance is. `desired_state` records
what you asked for: `running`, `stopped` or `deleted`. While a request is
being applied, the two can differ: a stop sets `desired_state=stopped` and
`current_state=stopping` until the guest powers off.

`current_state` also reports conditions observed on the guest:

* `crashed`: the guest crashed. If it is still desired to run, automatic
  recovery attempts a restart, at most once every 60 seconds. The crash is
  reported before recovery, so this state may be brief.
* `paused`: guest execution is paused.
* `suspended`: the guest entered power-management suspension.

Paused and suspended guests are not automatically restarted. You can stop a
crashed, paused or suspended instance, then start it again. Stopping also
cancels automatic crash recovery. For a crashed
guest, read the [console output](/compute/console) to investigate the cause,
even if automatic recovery has already brought it back to `running`.

## What each failure is telling you

<AccordionGroup>
  <Accordion title="current_state is error right after create" icon="triangle-alert">
    Read `faults`. Each entry carries a `code`, a `severity`, a `message` and
    `details` from the step that failed. More than one code can be active at
    once. See [Resource faults](/resource-faults) for the array itself.

    `POST /v1/instances/{instance_id}/start` accepts an instance in `error`, so
    a failure that happened while the guest was coming up can be retried
    without recreating anything. A start retry resolves `START_FAILED` and
    `LEGACY_FAILURE_STATE` only. It does not clear provisioning-owned codes.

    It does not rebuild what was never built. A failure earlier than that — the
    boot volume or the interfaces — leaves nothing for a start to bring up, and
    the instance is best deleted and created again once the cause is fixed. The
    usual causes are in the request: a subnet that does not route to an
    internet gateway while a NIC asked for a public address, a `user_data`
    document that is not valid YAML, or a boot size below the image's floor.

    | Code                     | Meaning and recovery                                                                                                      |
    | ------------------------ | ------------------------------------------------------------------------------------------------------------------------- |
    | `PROVISION_START_FAILED` | The provisioning saga did not start. Delete and create again once the cause is fixed. A start retry does not clear this.  |
    | `BOOT_VOLUME_FAILED`     | The boot disk was not built. Delete and create again. A start retry does not clear this.                                  |
    | `DATA_VOLUME_FAILED`     | A data volume was not attached during provisioning. Delete and create again. A start retry does not clear this.           |
    | `NIC_PROVISION_FAILED`   | A network interface was not provisioned. Delete and create again. A start retry does not clear this.                      |
    | `PUBLIC_IP_FAILED`       | A public address was not assigned. Delete and create again. A start retry does not clear this.                            |
    | `SEED_ISO_FAILED`        | The cloud-init seed was not built. Delete and create again. A start retry does not clear this.                            |
    | `VMSPEC_PUBLISH_FAILED`  | The instance spec was not published. Delete and create again. A start retry does not clear this.                          |
    | `START_FAILED`           | A start did not reach `running`. Retry start, or wait for reconcile.                                                      |
    | `STOP_FAILED`            | A stop did not reach `stopped`. Retry stop, or wait for reconcile.                                                        |
    | `REBOOT_FAILED`          | A reboot did not complete. Retry reboot, or wait for reconcile.                                                           |
    | `RECONCILE_FAILED`       | The host did not converge the current generation. A matching reconcile or power retry clears it.                          |
    | `MIGRATION_FAILED`       | A live move did not complete. This is a `warning` — the instance keeps serving on its source host.                        |
    | `RESIZE_PENDING_RESTART` | The flavor grew, but the running guest keeps its old size. This is a `warning` — stop and start the instance to apply it. |
  </Accordion>

  <Accordion title="The instance says running but nothing answers" icon="activity">
    Check `current_state`. A `running` guest can still have a network or
    guest-configuration problem; `running` does not mean SSH is ready.

    Read the [console output](/compute/console) next — it works on a guest that
    never reached the network, which is precisely the case SSH cannot
    diagnose.
  </Accordion>

  <Accordion title="My cloud-init never ran" icon="file-code">
    Two causes, both silent.

    Your `user_data` has to decode to a document starting with
    `#cloud-config`. A shell script is not merged and is not executed — put
    the commands in `runcmd`.

    Or the merge replaced something. Top-level keys replace wholesale, so a
    `users:` block of your own removes the default `basaltic` user and the
    authorized keys inside it, which reads from outside as "the instance
    booted and I cannot log in".
  </Accordion>

  <Accordion title="I cannot SSH in" icon="key-round">
    The default login user is **`basaltic`**, not `root` and not the image's
    own default — password login and root login are both disabled in the base
    configuration.

    Check that `keypairs` named a keypair that existed at launch, and that the
    NIC's security group admits your source on port 22. Confirm the address
    from `GET /v1/instances/{instance_id}/nics`: inspect each NIC's `addresses`
    and the nested `floating_ips`. The instance object has no IP summary fields.
  </Accordion>

  <Accordion title="Resize is refused with a capacity error" icon="cpu">
    An instance does not move hosts when it resizes, so the growth has to fit
    where it already is. A `dedicated` flavor needs whole free threads, not
    merely an undersubscribed host.

    The way around it is to launch a new instance on the target flavor — the
    scheduler is free to place that one anywhere — and move the workload, or
    to reinstall onto a fresh instance and reattach the data volumes.
  </Accordion>

  <Accordion title="A 409 on start, stop or reboot" icon="circle-x">
    Each action accepts a narrow set of states: start needs `stopped` or
    `error`; stop accepts `running`, `crashed`, `paused` or `suspended`;
    reboot needs `running`; resize and reinstall need
    `stopped`. Anything mid-transition — `pending`, `building`, `stopping`,
    `rebooting`, `deleting` — is refused until it settles. Poll `current_state` and
    retry.
  </Accordion>

  <Accordion title="An operation on an instance answers 404 but I can see it" icon="lock">
    Check `managed_by`. An instance owned by a load balancer or a database
    cluster is read-only through this API, and every mutation answers `404`
    rather than `403` — the customer surface does not admit it is actionable.

    A `403` is the other case, and a different one: the instance is yours and
    the IAM check on that specific action failed. Actions are per-operation, so
    a principal may read an instance and still not be allowed to stop it.
  </Accordion>
</AccordionGroup>

## Quotas

An instance holds three regional counters at once, all reserved together at
create and released at delete: `instances`, `vcpus` and `ram_mb`. That is why
a small `instances` limit is not the whole story — a handful of large flavors
can exhaust `vcpus` first.

`volumes_per_instance` caps the boot disk plus every data volume on one
instance, and is checked before the instance row exists, so exceeding it is a
plain `4xx` rather than an instance that provisions partway and parks in
`error`. Floating IPs come out of the network service's `floating_ips`.

<Warning>
  `instances` is rarely the limit you hit first. Three counters are reserved
  together at create and released together at delete, so a handful of large
  flavors exhausts `vcpus` or `ram_mb` while the instance count still looks
  comfortable. A quota refusal names which one ran out — read it rather than
  assuming.
</Warning>

Floating IPs come out of the network service's `floating_ips` allowance, which
[NAT gateways draw on too](/networking/troubleshooting#quotas).

## Next

<CardGroup cols={2}>
  <Card title="Console access" icon="terminal" href="/compute/console">
    The three ways in, including the one that works with no network.
  </Card>

  <Card title="Permissions" icon="key" href="/compute/permissions">
    When the refusal is a `403` rather than a state error.
  </Card>
</CardGroup>
