IDC business & buying guides

How automated server provisioning works: from paid order to delivered server

Automated server provisioning is an order state machine: a payment reserves a free server, then IP assignment, switch port setup, a PXE reinstall over IPMI and credential delivery run in a fixed order. Each step must detect failure and release what it took, or servers stay half-provisioned.

By Toplink product teamPublished 13 min read

Provisioning is a fixed-order pipeline

Automated server provisioning (also called hands-free delivery or dedicated server automation) solves exactly one problem: after the customer pays, the server gets an operating system, a working network connection and a set of credentials in the customer’s hands without anyone logging into an admin panel. It works by breaking delivery into steps with a fixed order, each with a defined input, output and success check, driven forward by an order state machine.

For a dedicated server, an order passes through these steps between payment and delivery:

Step Input Output Executed by Success check
1. Payment confirmation Payment gateway webhook Order marked paid Billing system Webhook signature valid, amount matches the order
2. Inventory matching Plan (site, carrier, CPU, memory, disk) One free server, reserved Data center management system A matching server with a reachable BMC was found and locked
3. IP assignment Site, carrier, number of IPs Primary IP, gateway, netmask, DNS; optional extra IPs and an IPv6 subnet IPAM module Unused, unlocked addresses taken from a block and marked in use
4. Switch port configuration The server’s switch port, VLAN, bandwidth Port up, correct VLAN, rate limit applied Switch management module Port state read back matches what was pushed
5. OS install Image, partition layout, root password, network parameters A bootable, reachable OS Install service + BMC Install callback received, SSH or RDP port answers
6. Credential delivery IP, password, hostname Customer notified, server visible in the portal Billing system Write-back succeeded, notification sent
7. Finishing Actual MAC address, port ARP binding in place, traffic metering started Switch and billing modules Binding applied, traffic pool counting

Every step needs a matching rollback action. Without one, servers get stuck in states like “IP assigned but no OS” or “port open but nobody using it”. The second half of this article covers how to design those rollbacks.

How cloud instances differ

A cloud instance never touches a BMC or a physical switch port, so the pipeline is much shorter: the billing system calls the virtualization platform’s API, clones a disk from a template, creates a virtual NIC on the right bridge or VLAN, injects the IP, password and SSH key through cloud-init metadata, and the instance is reachable within tens of seconds of booting.

The differences come down to three things:

Aspect Dedicated server Cloud instance
Inventory One physical machine per order, matched on hardware Scheduler picks a host with spare CPU, memory and storage
Network A real switch port that needs VLAN and rate-limit commands A virtual NIC controlled by security groups or flow rules, rate-limited on the host
Failure handling Step-by-step rollback; the machine may need a human look Destroy and recreate, at almost no cost

Because rebuilding a cloud instance is so cheap, many teams push dedicated server delivery in the same direction: wipe disks and reset the BMC automatically before every install, so that “run it again” is as reliable as possible.

The order state machine: states, events and idempotency

The state machine is the skeleton of the whole process. An order can only be in a small number of states, only specific events move it between them, and each transition triggers a fixed action:

Current state Event Next state Action
Pending payment Payment webhook Paid Create a provisioning task
Paid Task starts Provisioning Inventory match, IP, port, install
Provisioning All steps succeed Active Write back details, notify customer, start the billing cycle
Provisioning A step fails and retries are exhausted Failed Run compensating actions, alert operations
Active Invoice overdue Suspended Shut the switch port or power off
Suspended Renewal paid Active Reopen the port, power on
Active / Suspended Cancelled or expired Released Unbind the server, reclaim IPs, move the port to a quarantine VLAN

Two engineering rules are non-negotiable:

  • Events must be idempotent. Payment webhooks can arrive twice. Use the order number as a unique key so the second delivery returns success without creating a second task. The same goes for the install-complete callback.
  • Persist the state change before acting on it. Set the order to Provisioning and write the task record first, then touch the BMC and the switch. If the worker crashes halfway, the restarted worker can read the task record and see which step it reached.

Inventory matching: finding a free server safely

A plan is a set of conditions: site, carrier, CPU model or core count, memory, number and type of disks, bandwidth, number of IPs. Matching filters on those conditions and then applies three more checks:

  1. The server’s status is idle and no other running order has reserved it;
  2. The BMC address is reachable and the last power-state query succeeded;
  3. The automatically collected hardware inventory matches the plan. Hand-entered specs that differ from the real hardware are the main source of “this is not what I ordered” complaints after delivery.

Concurrency is the hard part: two orders paid in the same second must not get the same machine. The usual approach is a row lock in the database that finds and reserves in one transaction:

BEGIN;
SELECT id FROM servers
 WHERE datacenter_id = 3 AND plan_id = 17 AND status = 'idle'
 ORDER BY id LIMIT 1
 FOR UPDATE SKIP LOCKED;
-- suppose the query above returned id 1024
UPDATE servers SET status = 'reserved', order_id = 'ORD-20261003-0001' WHERE id = 1024;
COMMIT;

FOR UPDATE SKIP LOCKED makes concurrent transactions skip rows that are already locked, so each one gets a different server. When nothing is free, the order stays in Paid and operations get a restock alert; the customer does not get an error.

Network preparation: IP assignment and switch ports

IP assignment

Address blocks are registered in advance by site, carrier and type (public, IPMI, private), and each block is flagged as either available for automatic assignment or locked and reserved. Assignment takes the first free address from a matching block and records that block’s gateway, netmask and DNS servers, because those values go into the install answer file later. Extra IPs and an IPv6 subnet are taken at the same time according to the plan.

Assignment must be idempotent too: record “this order already holds these addresses” keyed by order number, so a retry reuses them instead of burning a few more IPs per failure. Before taking an address, probe it with ARP or ping; a reply means a leftover configuration still holds it, so skip it and flag the conflict.

Switch port configuration

Which switch and port a server is cabled to is recorded when it is racked. Provisioning pushes three things: VLAN, rate limit and port enable. The install phase adds a DHCP question: PXE needs DHCP, and production VLANs normally have no DHCP server. Two common designs:

Design How it works Advantage Cost
Install VLAN swap The port is moved to an install VLAN (with DHCP, TFTP and an HTTP image mirror) for the install, then back to the production VLAN No DHCP in the production network, image mirror isolated One extra VLAN change, and the install-complete check must be reliable
DHCP relay in the production VLAN DHCP relay on the production VLAN points to the install server; the answer file sets a static IP Port is configured once DHCP is exposed to customer segments and must answer only known MACs

On a Huawei S-series switch, putting the port into VLAN 100, limiting it to 100 Mbps in both directions (cir is in kbit/s) and enabling it looks like this. Some models do not support qos lr inbound and need qos car instead, so check the model’s manual. On Cisco IOS the VLAN is set with switchport access vlan 100 and the port enabled with no shutdown; rate limiting is done through a policy-map with police, whose syntax and supported directions vary by platform:

interface GigabitEthernet0/0/12
 port link-type access
 port default vlan 100
 qos lr inbound cir 100000
 qos lr outbound cir 100000
 undo shutdown

Always read the configuration back after pushing it: display interface GigabitEthernet0/0/12 for port state and display port vlan GigabitEthernet0/0/12 for VLAN membership on Huawei, or show interfaces GigabitEthernet1/0/12 switchport on Cisco. Only move on to the install when the readback matches. Pushing without reading back is the usual cause of “the port was configured but the server has no network”.

OS install and credential delivery: IPMI, PXE and answer files

Set the boot device and boot from PXE

Installing a dedicated server depends on the BMC: set the next boot to PXE, then reboot. With ipmitool that is two commands:

# Boot from PXE next time (add options=efiboot for UEFI machines)
ipmitool -I lanplus -H 10.20.0.15 -U admin -P 'password' chassis bootdev pxe options=efiboot
# Power cycle; the server comes up from PXE
ipmitool -I lanplus -H 10.20.0.15 -U admin -P 'password' chassis power cycle

On the PXE side, DHCP tells the NIC where to fetch the boot loader. A typical dnsmasq configuration hands out different boot files for BIOS and UEFI clients based on the client architecture option:

# /etc/dnsmasq.d/pxe.conf
dhcp-range=10.30.0.100,10.30.0.200,255.255.255.0,10m
# Client architectures 7 (EFI BC) and 9 (EFI x86-64) are both treated as UEFI
dhcp-match=set:efi,option:client-arch,7
dhcp-match=set:efi,option:client-arch,9
dhcp-boot=tag:efi,grubx64.efi
dhcp-boot=tag:!efi,pxelinux.0
enable-tftp
tftp-root=/srv/tftp

Answer files and the completion check

The answer file is what makes the install unattended: Kickstart for RHEL-family systems, preseed for Debian, autoinstall for Ubuntu, autounattend.xml for Windows. The provisioning service generates it per order, filling in the IP, gateway, DNS, hostname and password assigned earlier, and adds a callback that fires when the install finishes:

# Kickstart fragment, generated per order
network --bootproto=static --device=link --ip=203.0.113.37 --netmask=255.255.255.0 --gateway=203.0.113.1 --nameserver=203.0.113.53 --hostname=srv-0001
# Store only the SHA-512 hash of the password, never plain text
rootpw --iscrypted $6$<salt>$<hash>
# --nochroot: run the callback from the install environment, so it does not depend on curl in the new system
%post --nochroot
curl -fsS -X POST https://provision.example.com/callback/ORD-20261003-0001 -d status=installed
%end

Do not rely on a single signal to decide the install is done. A robust combination: receive the %post callback, then probe that SSH (Linux) or RDP (Windows) answers, then log in once with the delivered password and run hostname. Only when all three pass does the order move on. If no callback arrives before the deadline, grab a screenshot from the BMC console to see whether the server is stuck in a boot menu or on a disk error, then decide between a retry and a handoff to a person.

The alternative install path is mounting an ISO through the BMC’s virtual media and letting the answer file drive it, which suits small sites without a PXE environment. The drawback is that the image travels over the BMC’s network port, which is usually far slower than PXE over the production NIC.

Delivering credentials and finishing up

Once the OS is up, the provisioning service does four things:

  1. Write back: store the IP, hostname and initial password on the order so the customer portal can show the server. Show the password once at delivery and let the customer change it afterwards.
  2. Notify: send the delivery email or message without the plain-text password; point the customer to the portal instead.
  3. Bind: learn the NIC’s real MAC address from the switch and apply ARP binding and source guard so the IP cannot be spoofed. In the install-VLAN design, this is where the port moves back to the production VLAN.
  4. Start metering: add the port to its traffic pool and start billing by the plan’s bandwidth or traffic model. Starting the billing cycle at delivery rather than at payment is fairer, since the customer does not pay for the time spent installing.

The order is now Active. Reboots, reinstalls, rescue mode and password resets from here on are done by the customer in the self-service portal and never pass through the provisioning pipeline again.

Where provisioning fails and how to roll back

The hard part of automated server provisioning is not the happy path but the failure path. These are the common failure points in dedicated server delivery:

Step Typical failure How it is detected Rollback
Inventory matching No free server; server marked idle but BMC unreachable Empty query result; power-state query times out Keep the order in Paid and alert for restock; mark the unreachable server as needs-check and remove it from the pool
IP assignment Block exhausted; chosen address is live No free address; probe gets a reply Move to the next block; flag the live address as a conflict and pick another
Switch configuration SSH login fails; syntax differs on this model; readback mismatch Push times out or returns an error Retry once; if it still fails, put the port in a quarantine VLAN and shut it, release the IPs, mark the server needs-check
OS install PXE gets no DHCP lease; image mirror times out; old RAID metadata on disk; stuck in a boot menu No callback within the deadline; console screenshot Retry once automatically after wiping disk metadata; on a second failure hand off to a person and keep the IP and port configuration so they can continue
Write-back and notification Billing system API times out Non-2xx HTTP response Retry only the write-back; never rerun the install

A few principles for designing these rollbacks:

  • Compensate in reverse order. On an install failure, undo the port configuration first, then release the IPs, then update the server status. The other way round leaves a window where the IP is free but the port is still open.
  • Do not throw the server back into the pool blindly. A machine that failed to install may have a hardware fault; returning it to idle just makes the next order fail too. Mark it needs-check and let a person clear it.
  • Stop rolling back once the customer has the credentials. If a later step such as ARP binding fails after the write-back, alert but do not roll back, or you will take away a server the customer is already logging into.
  • Set a timeout on every step, but look before you fail it. Install time varies widely with image size, network speed and POST time. When the timeout fires, take a console screenshot and decide, rather than declaring failure outright.
  • Give every site redundant execution nodes. Installs and inventory collection should run on a node local to each data center, and when one node is down another must take over, or provisioning for that whole site stops.

How this looks in a management system

The pipeline above needs the billing system and the data center management system to cooperate. In Toplink’s setup, CloudFinance drives provisioning, suspension and release with an order state machine, and after payment a dedicated provisioning module syncs the order to Toplink DCIM. DCIM matches a free server on the plan’s CPU, memory, disk, IP and bandwidth, takes addresses from blocks with auto-assign enabled, pushes the port configuration through the vendor’s command set, and reinstalls the OS over IPMI from a catalog that includes CentOS, AlmaLinux, Rocky, Ubuntu, Debian, Windows and ESXi. Agent nodes in each site run reinstalls and hardware collection locally, another node is picked automatically if one is unavailable, and install progress is visible in the task queue. See billing and automatic provisioning, IP address management and IPMI remote management for the details. Providers already running WHMCS can use the integration plugin to get the same chain: provision on payment, write back the IP and password, power off and disconnect when overdue, restore on renewal, and unbind the server and release its IPs on cancellation.

FAQ

How long does automated provisioning of a dedicated server take?

From payment to the start of the install is usually seconds to a few minutes. Most of the time goes into the OS install itself, which depends on image size, whether the image mirror sits inside the data center and how long the server's POST takes; Windows is normally slower than Linux. A local mirror and wiping old RAID metadata before the install shorten it noticeably.

Can a server without IPMI be provisioned automatically?

It can be installed, but not hands-free. Without out-of-band management there is no way to set the boot device and reboot remotely, and no screen to look at when the install fails, so someone has to go to the rack. Every server in an automated pool should have a BMC on the management network.

How is cloud instance provisioning different from dedicated servers?

A cloud instance is cloned from a template through the virtualization platform's API, gets its network settings and password from cloud-init and is ready in tens of seconds; on failure it is simply destroyed and recreated. A dedicated server involves a real BMC, a real switch port and a full OS install, so it takes longer, fails in more places and must release its IP and port configuration on rollback.

What should the customer see when provisioning fails?

The order should stay in a provisioning state and alert operations immediately, rather than showing the customer a raw error. If the order is still not delivered after a set deadline, hand it to a person and tell the customer the expected delivery time.

See how it works in your data center.

Start with a product demo and map out your next step.

View pricing
Hotline 400-112-2951

Let’s talk about your IDC

Scan with WeChat to book a product demo or request a trial.

Toplink WeCom contact QR code

Save this code or scan it with WeChat / WeCom

Or call
400-112-2951
Telegram
@TopLink88