How automated server provisioning works: from paid order to delivered server
Automated server provisioning is an order state machine: a payment reserves a free server, then IP assignment, switch port setup, a PXE reinstall over IPMI and credential delivery run in a fixed order. Each step must detect failure and release what it took, or servers stay half-provisioned.
Provisioning is a fixed-order pipeline
Automated server provisioning (also called hands-free delivery or dedicated server automation) solves exactly one problem: after the customer pays, the server gets an operating system, a working network connection and a set of credentials in the customer’s hands without anyone logging into an admin panel. It works by breaking delivery into steps with a fixed order, each with a defined input, output and success check, driven forward by an order state machine.
For a dedicated server, an order passes through these steps between payment and delivery:
| Step | Input | Output | Executed by | Success check |
|---|---|---|---|---|
| 1. Payment confirmation | Payment gateway webhook | Order marked paid | Billing system | Webhook signature valid, amount matches the order |
| 2. Inventory matching | Plan (site, carrier, CPU, memory, disk) | One free server, reserved | Data center management system | A matching server with a reachable BMC was found and locked |
| 3. IP assignment | Site, carrier, number of IPs | Primary IP, gateway, netmask, DNS; optional extra IPs and an IPv6 subnet | IPAM module | Unused, unlocked addresses taken from a block and marked in use |
| 4. Switch port configuration | The server’s switch port, VLAN, bandwidth | Port up, correct VLAN, rate limit applied | Switch management module | Port state read back matches what was pushed |
| 5. OS install | Image, partition layout, root password, network parameters | A bootable, reachable OS | Install service + BMC | Install callback received, SSH or RDP port answers |
| 6. Credential delivery | IP, password, hostname | Customer notified, server visible in the portal | Billing system | Write-back succeeded, notification sent |
| 7. Finishing | Actual MAC address, port | ARP binding in place, traffic metering started | Switch and billing modules | Binding applied, traffic pool counting |
Every step needs a matching rollback action. Without one, servers get stuck in states like “IP assigned but no OS” or “port open but nobody using it”. The second half of this article covers how to design those rollbacks.
How cloud instances differ
A cloud instance never touches a BMC or a physical switch port, so the pipeline is much shorter: the billing system calls the virtualization platform’s API, clones a disk from a template, creates a virtual NIC on the right bridge or VLAN, injects the IP, password and SSH key through cloud-init metadata, and the instance is reachable within tens of seconds of booting.
The differences come down to three things:
| Aspect | Dedicated server | Cloud instance |
|---|---|---|
| Inventory | One physical machine per order, matched on hardware | Scheduler picks a host with spare CPU, memory and storage |
| Network | A real switch port that needs VLAN and rate-limit commands | A virtual NIC controlled by security groups or flow rules, rate-limited on the host |
| Failure handling | Step-by-step rollback; the machine may need a human look | Destroy and recreate, at almost no cost |
Because rebuilding a cloud instance is so cheap, many teams push dedicated server delivery in the same direction: wipe disks and reset the BMC automatically before every install, so that “run it again” is as reliable as possible.
The order state machine: states, events and idempotency
The state machine is the skeleton of the whole process. An order can only be in a small number of states, only specific events move it between them, and each transition triggers a fixed action:
| Current state | Event | Next state | Action |
|---|---|---|---|
| Pending payment | Payment webhook | Paid | Create a provisioning task |
| Paid | Task starts | Provisioning | Inventory match, IP, port, install |
| Provisioning | All steps succeed | Active | Write back details, notify customer, start the billing cycle |
| Provisioning | A step fails and retries are exhausted | Failed | Run compensating actions, alert operations |
| Active | Invoice overdue | Suspended | Shut the switch port or power off |
| Suspended | Renewal paid | Active | Reopen the port, power on |
| Active / Suspended | Cancelled or expired | Released | Unbind the server, reclaim IPs, move the port to a quarantine VLAN |
Two engineering rules are non-negotiable:
- Events must be idempotent. Payment webhooks can arrive twice. Use the order number as a unique key so the second delivery returns success without creating a second task. The same goes for the install-complete callback.
- Persist the state change before acting on it. Set the order to Provisioning and write the task record first, then touch the BMC and the switch. If the worker crashes halfway, the restarted worker can read the task record and see which step it reached.
Inventory matching: finding a free server safely
A plan is a set of conditions: site, carrier, CPU model or core count, memory, number and type of disks, bandwidth, number of IPs. Matching filters on those conditions and then applies three more checks:
- The server’s status is idle and no other running order has reserved it;
- The BMC address is reachable and the last power-state query succeeded;
- The automatically collected hardware inventory matches the plan. Hand-entered specs that differ from the real hardware are the main source of “this is not what I ordered” complaints after delivery.
Concurrency is the hard part: two orders paid in the same second must not get the same machine. The usual approach is a row lock in the database that finds and reserves in one transaction:
BEGIN;
SELECT id FROM servers
WHERE datacenter_id = 3 AND plan_id = 17 AND status = 'idle'
ORDER BY id LIMIT 1
FOR UPDATE SKIP LOCKED;
-- suppose the query above returned id 1024
UPDATE servers SET status = 'reserved', order_id = 'ORD-20261003-0001' WHERE id = 1024;
COMMIT;
FOR UPDATE SKIP LOCKED makes concurrent transactions skip rows that are already locked, so each one gets a different server. When nothing is free, the order stays in Paid and operations get a restock alert; the customer does not get an error.
Network preparation: IP assignment and switch ports
IP assignment
Address blocks are registered in advance by site, carrier and type (public, IPMI, private), and each block is flagged as either available for automatic assignment or locked and reserved. Assignment takes the first free address from a matching block and records that block’s gateway, netmask and DNS servers, because those values go into the install answer file later. Extra IPs and an IPv6 subnet are taken at the same time according to the plan.
Assignment must be idempotent too: record “this order already holds these addresses” keyed by order number, so a retry reuses them instead of burning a few more IPs per failure. Before taking an address, probe it with ARP or ping; a reply means a leftover configuration still holds it, so skip it and flag the conflict.
Switch port configuration
Which switch and port a server is cabled to is recorded when it is racked. Provisioning pushes three things: VLAN, rate limit and port enable. The install phase adds a DHCP question: PXE needs DHCP, and production VLANs normally have no DHCP server. Two common designs:
| Design | How it works | Advantage | Cost |
|---|---|---|---|
| Install VLAN swap | The port is moved to an install VLAN (with DHCP, TFTP and an HTTP image mirror) for the install, then back to the production VLAN | No DHCP in the production network, image mirror isolated | One extra VLAN change, and the install-complete check must be reliable |
| DHCP relay in the production VLAN | DHCP relay on the production VLAN points to the install server; the answer file sets a static IP | Port is configured once | DHCP is exposed to customer segments and must answer only known MACs |
On a Huawei S-series switch, putting the port into VLAN 100, limiting it to 100 Mbps in both directions (cir is in kbit/s) and enabling it looks like this. Some models do not support qos lr inbound and need qos car instead, so check the model’s manual. On Cisco IOS the VLAN is set with switchport access vlan 100 and the port enabled with no shutdown; rate limiting is done through a policy-map with police, whose syntax and supported directions vary by platform:
interface GigabitEthernet0/0/12
port link-type access
port default vlan 100
qos lr inbound cir 100000
qos lr outbound cir 100000
undo shutdown
Always read the configuration back after pushing it: display interface GigabitEthernet0/0/12 for port state and display port vlan GigabitEthernet0/0/12 for VLAN membership on Huawei, or show interfaces GigabitEthernet1/0/12 switchport on Cisco. Only move on to the install when the readback matches. Pushing without reading back is the usual cause of “the port was configured but the server has no network”.
OS install and credential delivery: IPMI, PXE and answer files
Set the boot device and boot from PXE
Installing a dedicated server depends on the BMC: set the next boot to PXE, then reboot. With ipmitool that is two commands:
# Boot from PXE next time (add options=efiboot for UEFI machines)
ipmitool -I lanplus -H 10.20.0.15 -U admin -P 'password' chassis bootdev pxe options=efiboot
# Power cycle; the server comes up from PXE
ipmitool -I lanplus -H 10.20.0.15 -U admin -P 'password' chassis power cycle
On the PXE side, DHCP tells the NIC where to fetch the boot loader. A typical dnsmasq configuration hands out different boot files for BIOS and UEFI clients based on the client architecture option:
# /etc/dnsmasq.d/pxe.conf
dhcp-range=10.30.0.100,10.30.0.200,255.255.255.0,10m
# Client architectures 7 (EFI BC) and 9 (EFI x86-64) are both treated as UEFI
dhcp-match=set:efi,option:client-arch,7
dhcp-match=set:efi,option:client-arch,9
dhcp-boot=tag:efi,grubx64.efi
dhcp-boot=tag:!efi,pxelinux.0
enable-tftp
tftp-root=/srv/tftp
Answer files and the completion check
The answer file is what makes the install unattended: Kickstart for RHEL-family systems, preseed for Debian, autoinstall for Ubuntu, autounattend.xml for Windows. The provisioning service generates it per order, filling in the IP, gateway, DNS, hostname and password assigned earlier, and adds a callback that fires when the install finishes:
# Kickstart fragment, generated per order
network --bootproto=static --device=link --ip=203.0.113.37 --netmask=255.255.255.0 --gateway=203.0.113.1 --nameserver=203.0.113.53 --hostname=srv-0001
# Store only the SHA-512 hash of the password, never plain text
rootpw --iscrypted $6$<salt>$<hash>
# --nochroot: run the callback from the install environment, so it does not depend on curl in the new system
%post --nochroot
curl -fsS -X POST https://provision.example.com/callback/ORD-20261003-0001 -d status=installed
%end
Do not rely on a single signal to decide the install is done. A robust combination: receive the %post callback, then probe that SSH (Linux) or RDP (Windows) answers, then log in once with the delivered password and run hostname. Only when all three pass does the order move on. If no callback arrives before the deadline, grab a screenshot from the BMC console to see whether the server is stuck in a boot menu or on a disk error, then decide between a retry and a handoff to a person.
The alternative install path is mounting an ISO through the BMC’s virtual media and letting the answer file drive it, which suits small sites without a PXE environment. The drawback is that the image travels over the BMC’s network port, which is usually far slower than PXE over the production NIC.
Delivering credentials and finishing up
Once the OS is up, the provisioning service does four things:
- Write back: store the IP, hostname and initial password on the order so the customer portal can show the server. Show the password once at delivery and let the customer change it afterwards.
- Notify: send the delivery email or message without the plain-text password; point the customer to the portal instead.
- Bind: learn the NIC’s real MAC address from the switch and apply ARP binding and source guard so the IP cannot be spoofed. In the install-VLAN design, this is where the port moves back to the production VLAN.
- Start metering: add the port to its traffic pool and start billing by the plan’s bandwidth or traffic model. Starting the billing cycle at delivery rather than at payment is fairer, since the customer does not pay for the time spent installing.
The order is now Active. Reboots, reinstalls, rescue mode and password resets from here on are done by the customer in the self-service portal and never pass through the provisioning pipeline again.
Where provisioning fails and how to roll back
The hard part of automated server provisioning is not the happy path but the failure path. These are the common failure points in dedicated server delivery:
| Step | Typical failure | How it is detected | Rollback |
|---|---|---|---|
| Inventory matching | No free server; server marked idle but BMC unreachable | Empty query result; power-state query times out | Keep the order in Paid and alert for restock; mark the unreachable server as needs-check and remove it from the pool |
| IP assignment | Block exhausted; chosen address is live | No free address; probe gets a reply | Move to the next block; flag the live address as a conflict and pick another |
| Switch configuration | SSH login fails; syntax differs on this model; readback mismatch | Push times out or returns an error | Retry once; if it still fails, put the port in a quarantine VLAN and shut it, release the IPs, mark the server needs-check |
| OS install | PXE gets no DHCP lease; image mirror times out; old RAID metadata on disk; stuck in a boot menu | No callback within the deadline; console screenshot | Retry once automatically after wiping disk metadata; on a second failure hand off to a person and keep the IP and port configuration so they can continue |
| Write-back and notification | Billing system API times out | Non-2xx HTTP response | Retry only the write-back; never rerun the install |
A few principles for designing these rollbacks:
- Compensate in reverse order. On an install failure, undo the port configuration first, then release the IPs, then update the server status. The other way round leaves a window where the IP is free but the port is still open.
- Do not throw the server back into the pool blindly. A machine that failed to install may have a hardware fault; returning it to idle just makes the next order fail too. Mark it needs-check and let a person clear it.
- Stop rolling back once the customer has the credentials. If a later step such as ARP binding fails after the write-back, alert but do not roll back, or you will take away a server the customer is already logging into.
- Set a timeout on every step, but look before you fail it. Install time varies widely with image size, network speed and POST time. When the timeout fires, take a console screenshot and decide, rather than declaring failure outright.
- Give every site redundant execution nodes. Installs and inventory collection should run on a node local to each data center, and when one node is down another must take over, or provisioning for that whole site stops.
How this looks in a management system
The pipeline above needs the billing system and the data center management system to cooperate. In Toplink’s setup, CloudFinance drives provisioning, suspension and release with an order state machine, and after payment a dedicated provisioning module syncs the order to Toplink DCIM. DCIM matches a free server on the plan’s CPU, memory, disk, IP and bandwidth, takes addresses from blocks with auto-assign enabled, pushes the port configuration through the vendor’s command set, and reinstalls the OS over IPMI from a catalog that includes CentOS, AlmaLinux, Rocky, Ubuntu, Debian, Windows and ESXi. Agent nodes in each site run reinstalls and hardware collection locally, another node is picked automatically if one is unavailable, and install progress is visible in the task queue. See billing and automatic provisioning, IP address management and IPMI remote management for the details. Providers already running WHMCS can use the integration plugin to get the same chain: provision on payment, write back the IP and password, power off and disconnect when overdue, restore on renewal, and unbind the server and release its IPs on cancellation.
FAQ
How long does automated provisioning of a dedicated server take?
From payment to the start of the install is usually seconds to a few minutes. Most of the time goes into the OS install itself, which depends on image size, whether the image mirror sits inside the data center and how long the server's POST takes; Windows is normally slower than Linux. A local mirror and wiping old RAID metadata before the install shorten it noticeably.
Can a server without IPMI be provisioned automatically?
It can be installed, but not hands-free. Without out-of-band management there is no way to set the boot device and reboot remotely, and no screen to look at when the install fails, so someone has to go to the rack. Every server in an automated pool should have a BMC on the management network.
How is cloud instance provisioning different from dedicated servers?
A cloud instance is cloned from a template through the virtualization platform's API, gets its network settings and password from cloud-init and is ready in tens of seconds; on failure it is simply destroyed and recreated. A dedicated server involves a real BMC, a real switch port and a full OS install, so it takes longer, fails in more places and must release its IP and port configuration on rollback.
What should the customer see when provisioning fails?
The order should stay in a provisioning state and alert operations immediately, rather than showing the customer a raw error. If the order is still not delivered after a set deadline, hand it to a person and tell the customer the expected delivery time.