Technology

Private Inference, Operated as One Stack

LLM Machines configures and maintains the infrastructure layer, inference runtime, gateway and policy, operational Console, Observability layer, and lifecycle and recovery controls. Your favourite apps and harnesses connect through a documented API handoff. Nothing leaves your infrastructure.

01 · Architecture

Reference architecture.

A managed inference core runs inside your infrastructure. Customer applications connect only through the documented Product edge and remain under customer ownership.

Customer environment
Teams & Users

Developers · analysts · support · legal · operations.
Use the client experiences they already know.

Authenticate through enterprise identity
Customer environment
Enterprise Identity Provider

Federate the identity provider your organisation already operates.

OIDC · SAML 2.0
Customer application boundary
Favourite Apps & Harnesses

Connect client applications and developer tools to approved local models.

Models · Chat Completions
LLM Machines · Customer-Premises Inference Stack Managed operational boundary

Infrastructure Layer

Compute · memory · storage · network · operating platform · physical controls

Workload sizing · distributor-backed Supermicro procurement · customer-controlled environment

CORE

Local Inference Runtime

Qualified open-weight model serving configured for throughput, latency, quality and capacity targets

Approved models · local model registry · managed configuration

CORE

Gateway & Policy

/models · streaming and non-streaming /chat/completions · credentials · model access · routing · rate policy

Enterprise identity federation · documented API handoff

CORE

Operational Console

CORE
Overview
system status · alerts
Operational summary
Model Estate
models · routes · capacity
Inference control
Hardware
capacity · health · inventory
Appliance status
Access & Activity
credentials · metadata · events
Policy and evidence

Observability Layer

Model health · capacity · operational metrics · alerts · hardware telemetry

Local operational visibility · capacity reviews · support evidence

CORE

Lifecycle & Recovery

Controlled updates · preflight · backup · restore · rollback · recovery history

Managed maintenance · documented recovery procedures

CORE

Product Edge

Sole customer ingress · authentication boundary · no-bypass enforcement · approved service routing

Documented Application API · private administrative routes

CORE
CUSTOMER APPLICATION BOUNDARY · MANAGED INFERENCE OPERATIONS REMAIN AVAILABLE

Customer Applications

Your favourite client applications, developer tools and harnesses connect to approved local models

Each application receives its own credential, approved model access and policy controls through the documented local interface.

API handoff · application credentials · model policy · customer-owned experience

CLIENT
Managed inference

One operational boundary.

LLM Machines configures and maintains the infrastructure, inference, policy, Console, Observability, update and recovery layers as one managed stack.

API handoff

Connect the tools your teams prefer.

Use one documented contract across the client applications and harnesses your teams already prefer.

Source rights

Inspect the first-party Product layer.

Original first-party LLM Machines Product source is source-available under the unmodified PolyForm Internal Use License 1.0.0. Third-party components retain their upstream licences.

02 · API handoff

Connect applications through a stable contract.

Engineering teams receive a documented local interface, approved model access and application-specific credentials at handoff.

Supported API

Point your client to local inference.

Chat Completions is the documented API protocol your favourite apps and harnesses use to access approved local models. Use /models to discover the approved model set.

Policy

Handoff with clear controls.

Application credentials, model access, routing and rate policies are configured against the customer application boundary.

Customer boundary

Keep applications in customer ownership.

Application selection, permissions, data and workflows stay outside the standard Product. Optional integration consulting uses a separate scope and documented handoff.

03 · FAQ

Architecture questions.

Details security, platform and developer teams usually ask before approving an on-prem AI deployment.

Can existing API clients use LLM Machines?

Yes. Existing clients can connect through the documented Chat Completions protocol, while the /models endpoint exposes the approved model set.

Can the appliance operate without public-internet access?

Yes. Signed offline release and update bundles plus local model registries support isolated operation inside your infrastructure.

Which models can run locally?

The architecture is designed for open-weight model families such as Llama, Mistral and Qwen, with model choice sized to your hardware, latency and quality requirements.

How do applications authenticate?

Each Application receives its own credential, model access and operational controls through the local Console.

How is Product source licensed?

Original first-party LLM Machines Product source is source-available under the unmodified PolyForm Internal Use License 1.0.0. Third-party components retain their upstream licences.

What is inside the managed inference core?

The managed inference core includes the infrastructure platform, local inference runtime, gateway and policy, operational Console, identity, Observability layer, and lifecycle and recovery controls.

What's next

Ready to dig deeper?

See how the technology lands inside your environment through onboarding and a defined service scope, or talk to us.