On-prem AI appliance

Managed private inference, inside your infrastructure.

LLM Machines sizes, configures, monitors and maintains private inference inside infrastructure you control. Connect your favourite apps and harnesses through a documented API while keeping the applications, permissions, data and workflows under your ownership. Nothing leaves your infrastructure.

01 · Managed inference

From workload sizing to operational handoff.

The appliance combines distributor-backed infrastructure with a configured inference environment and a recurring operating service.

Sizing & procurement

Infrastructure matched to the workload.

Size GPU, node, storage and network requirements, then coordinate distributor-backed Supermicro procurement after customer approval.

Inference configuration

Approved models, configured locally.

Configure model serving, routes, credentials, quotas and access policy for the latency, quality and throughput you need.

Observability

Operational visibility.

Monitor model health, capacity, request metadata, alerts and hardware through the local Observability layer.

Updates & recovery

A maintained operating baseline.

Manage controlled updates, backup, recovery and isolation procedures through agreed maintenance windows.

API handoff

Connect your favourite apps and harnesses.

Chat Completions is the documented API protocol your favourite apps and harnesses use to access approved local models.

Source rights

Use and modify first-party Product source internally.

Original first-party LLM Machines Product source is source-available under the unmodified PolyForm Internal Use License 1.0.0. Third-party components retain their upstream licences.

02 · Responsibility boundary

Designed for customer-controlled infrastructure.

Inference, operational controls and approved advanced-integration boundaries run inside the environment you control.

For IT and security teams.

Identity federation, Admin and Operator roles, network pre-flight, metadata-only audit and signed lifecycle controls are handled during onboarding.

  • OIDC / SAML identity federation
  • Signed offline release verification
  • Local Application credentials

For application teams.

LLM Machines provides local inference and a documented API handoff. Your teams retain their chosen applications, permissions, data and workflow ownership.

  • Favourite apps and harnesses
  • Separately scoped advanced integrations
  • Customer-controlled operations
Next

Size an appliance for your workload.

Review the architecture, deployment plan and pricing model before a discovery call.