Point your client to local inference.
Chat Completions is the documented API protocol your favourite apps and harnesses use to access approved local models. Use /models to discover the approved model set.
LLM Machines configures and maintains the infrastructure layer, inference runtime, gateway and policy, operational Console, Observability layer, and lifecycle and recovery controls. Your favourite apps and harnesses connect through a documented API handoff. Nothing leaves your infrastructure.
A managed inference core runs inside your infrastructure. Customer applications connect only through the documented Product edge and remain under customer ownership.
Developers · analysts · support · legal · operations.
Use the client experiences they already know.
Federate the identity provider your organisation already operates.
Connect client applications and developer tools to approved local models.
Compute · memory · storage · network · operating platform · physical controls
Workload sizing · distributor-backed Supermicro procurement · customer-controlled environment
Qualified open-weight model serving configured for throughput, latency, quality and capacity targets
Approved models · local model registry · managed configuration
/models · streaming and non-streaming /chat/completions · credentials · model access · routing · rate policy
Enterprise identity federation · documented API handoff
Model health · capacity · operational metrics · alerts · hardware telemetry
Local operational visibility · capacity reviews · support evidence
Controlled updates · preflight · backup · restore · rollback · recovery history
Managed maintenance · documented recovery procedures
Sole customer ingress · authentication boundary · no-bypass enforcement · approved service routing
Documented Application API · private administrative routes
Your favourite client applications, developer tools and harnesses connect to approved local models
Each application receives its own credential, approved model access and policy controls through the documented local interface.
API handoff · application credentials · model policy · customer-owned experience
LLM Machines configures and maintains the infrastructure, inference, policy, Console, Observability, update and recovery layers as one managed stack.
Use one documented contract across the client applications and harnesses your teams already prefer.
Original first-party LLM Machines Product source is source-available under the unmodified PolyForm Internal Use License 1.0.0. Third-party components retain their upstream licences.
Engineering teams receive a documented local interface, approved model access and application-specific credentials at handoff.
Chat Completions is the documented API protocol your favourite apps and harnesses use to access approved local models. Use /models to discover the approved model set.
Application credentials, model access, routing and rate policies are configured against the customer application boundary.
Application selection, permissions, data and workflows stay outside the standard Product. Optional integration consulting uses a separate scope and documented handoff.
Details security, platform and developer teams usually ask before approving an on-prem AI deployment.
Yes. Existing clients can connect through the documented Chat Completions protocol, while the /models endpoint exposes the approved model set.
Yes. Signed offline release and update bundles plus local model registries support isolated operation inside your infrastructure.
The architecture is designed for open-weight model families such as Llama, Mistral and Qwen, with model choice sized to your hardware, latency and quality requirements.
Each Application receives its own credential, model access and operational controls through the local Console.
Original first-party LLM Machines Product source is source-available under the unmodified PolyForm Internal Use License 1.0.0. Third-party components retain their upstream licences.
The managed inference core includes the infrastructure platform, local inference runtime, gateway and policy, operational Console, identity, Observability layer, and lifecycle and recovery controls.
See how the technology lands inside your environment through onboarding and a defined service scope, or talk to us.