Onboarding

Deploy On-Prem AI with a Clear Handoff

A tailored deployment, not a SaaS sign-up. We size and configure private inference, validate the operational controls and hand over a documented API for the applications your team already owns.

00 · Deployment path

From discovery to handoff.

Standard onboarding ends with an accepted inference environment, a validated API handoff and a documented maintenance plan.

01

Discovery and sizing.

Map workloads, models, concurrency, latency, availability and growth before the infrastructure is finalised.

02

Distributor-backed procurement.

Translate the approved design into enterprise Supermicro infrastructure fulfilled through a distributor-backed channel.

03

Infrastructure readiness.

Validate rack, power, cooling, network, identity, certificates, storage and the agreed support-access posture.

04

Inference configuration.

Configure approved models, routes, policies, credentials, observability, updates, backup and recovery.

05

Operational validation and API handoff.

Benchmark the supported routes and validate the documented API with one representative customer-owned application.

06

Handoff and maintenance.

Deliver the runbook, train administrators and begin the agreed monitoring, update and capacity-review cadence.

01 · The operating model

One supported core. Your applications stay familiar.

The supported core provides managed private inference and a documented API contract. Your applications, permissions, data and workflows remain under your control and outside the standard Product.

Core Appliance

Local inference and operational control.

Model serving, inference gateway, Console, identity, observability layer, metadata-only audit, offline updates, backup and recovery.

API handoff

Connect your favourite apps and harnesses.

Chat Completions is the documented API protocol your applications use to access approved local models while the application experience stays under your control.

Source and maintenance

Delivered rights stay separate from maintenance.

Delivered original first-party Product source is source-available and can be used and changed internally under the unmodified PolyForm Internal Use License 1.0.0. Third-party components retain their upstream licences. Maintenance controls future qualified releases, recovery retrieval and support, not the current runtime.

Emergency Isolation

A controlled Admin action pauses Application inference at the Product edge while keeping the Console, identity, audit, hardware monitoring, backup and recovery available.

02 · The deployment sequence

Six stages. One clear acceptance boundary.

Each stage has explicit owners and exit criteria. Standard delivery concludes with operational acceptance and a documented API handoff.

00 Discovery

Discovery and sizing.

Document workload patterns, model requirements, concurrency, latency, availability, data sensitivity and capacity growth. The result is an approved infrastructure design and a clear responsibility boundary.

Tailored to: your workloads, models and operating requirements
01 Procurement

Distributor-backed procurement.

LLM Machines translates the approved design into a bill of materials and coordinates enterprise Supermicro fulfilment through a distributor-backed channel. Availability, lead time and applicable manufacturer or distributor warranty terms are documented before ordering.

Accepted when: configuration, fulfilment and warranty path are approved
02 Readiness

Infrastructure readiness.

Rack, power, cooling, network, DNS, NTP, certificates, identity, storage, backup targets and the approved support-access posture are validated. Hardware health and throughput tests must pass before inference configuration begins.

Tailored to: your site, security and network posture
03 Configuration

Inference configuration.

Approved models, routes, aliases, context limits, credentials, quotas and access policies are configured alongside the Console, identity, observability, update, backup, recovery and isolation controls.

Tailored to: your approved models, policies and capacity limits
04 Acceptance

Operational validation and API handoff.

Benchmark the supported model routes and validate credentials, quotas, observability, metadata-only audit, updates, backup, recovery and isolation. One customer-owned application completes an API smoke test using the documented Models and Chat Completions endpoints.

Accepted when: operational controls and the representative API request pass
05 Operations

Handoff and maintenance.

Administrators receive the as-built configuration, runbooks and training for routes, credentials, alerts, updates, backup and recovery. The agreed monitoring, maintenance and capacity-review cadence begins after sign-off.

Tailored to: your operating windows, support access and review cadence
03 · Who does what

RACI summary.

Explicit ownership for every onboarding activity. R = Responsible (does the work) · A = Accountable (signs off) · C = Consulted · I = Informed.

Activity Us Client IT Application Owner Operator
Pre-flight & questionnaireA / RRCI
Hardware rack & powerCR / AII
First boot & release verificationR / ACII
Network, TLS, DNSCR / AII
IdP federationRA / RII
Inference core configurationR / ACIC
Application credentials & model scopeRCA / RC
API handoff smoke testRCA / RC
Admin trainingR / AAII
Customer application operationCIA / RR
Optional application integrationRCAC
Day-2 operationsCA / RCR
04 · FAQ

Deployment questions.

What security, IT and business teams usually need to know before scheduling an on-prem AI rollout.

What do we need before deployment starts?

You need workload and capacity requirements, infrastructure constraints, identity details, approved models, application owners and representative API requests for validation.

Can onboarding work for isolated environments?

Yes. Signed offline release bundles, local model bundles and signed manual update bundles support deployments without a public-internet dependency.

Which identity providers are supported?

The appliance supports standard enterprise identity federation through OIDC or SAML, with role mapping for Admin and Operator access.

How do applications connect?

Your chosen applications connect through the documented Chat Completions protocol, while your team keeps control of the application experience and workflow.

What makes a good pilot workload?

A good pilot has real users, clear success criteria, representative workload patterns and enough operational depth to validate the full deployment.

What's next

Pilot the appliance against your real workloads.

Discovery, a sized infrastructure design, controlled deployment, API acceptance and ongoing maintenance from one accountable team.