LLM Machines · Part 1 of 2
A maintenance manual can contain a customer's site layout, equipment serial numbers and years of engineering changes. Give an AI application access to that material and the infrastructure decision becomes quite specific. The business needs to know where the document is processed, which systems retain a copy, and who has permission to inspect it. Once service engineers depend on the application during a working shift, its availability matters too.
LLM Machines configures and maintains private AI inference environments for businesses that need to make those decisions themselves. The environment runs at a customer site or in a data centre the customer contracts for and controls. We handle the agreed configuration and operating work, including model serving, maintenance and capacity reviews.
Inference is the work a model does when it receives an input and produces an output. It sits underneath the application people use. An engineering assistant, for example, might gather relevant passages from a manual and send them to a model. Our service covers the inference environment behind that application. The business or its software supplier remains responsible for the application, its documents and its users.
The number of businesses using AI has grown. Eurostat reported that 20% of EU enterprises with at least ten employees used AI technologies in 2025, up from 13.5% in 2024. The survey covers several forms of AI across its included business sectors, well beyond language models. We read those figures as a reason to expect more AI-related decisions on an IT team's desk, alongside the systems it already maintains. Eurostat, December 2025.
The suppliers are making deployment location a more explicit choice as well. On 11 August 2026, Mistral announced generally available regional inference endpoints, allowing customers to choose processing in Europe or the United States. Its announcement also describes limited, safeguarded transfers to subprocessors outside the selected region. That qualification belongs in a buyer's review alongside the regional label: understanding the complete processing arrangement requires more than reading the location of the model server. Mistral's regional inference announcement.
For European organisations, there is also a recent regulatory development. On 2 August 2026, the European Commission's AI Office and national authorities took on enforcement powers under the AI Act, while new transparency requirements began applying to certain AI systems. The obligations depend on the system and the organisation's role. A private server does not, by itself, resolve them. We see the practical consequence for a buyer as a need to assign responsibility for the system's use and to keep records of the decisions made around it. European Commission announcement, July 2026.
For a company planning to use AI over several years, the purchase extends beyond the model available today. The proposal needs to explain how the service will accommodate a change in demand, how a model replacement will be evaluated and who can authorise it. These are decisions a business can put into an operating agreement before it becomes dependent on the service.
A hosted inference API can be a sensible starting point. The provider operates the serving infrastructure, and the customer can begin without buying a server. The terms for model availability, processing, usage limits and support become part of the selection. For a workload used occasionally, that arrangement may be easier to justify than maintaining dedicated capacity.
Running a model within infrastructure the customer controls gives the customer more say over the installed configuration and change schedule. That can matter when a business needs to retain a tested model version, restrict support access or fit maintenance around its own operations. It also brings work: someone must monitor the runtime, keep its dependencies maintained and restore the service after a failure. A business with the people and capacity to do that internally may choose to own the whole operation.
The financial comparison needs similar care. A dedicated server carries costs while it is idle, and its useful capacity depends on the requests being sent to it. A short extraction task and a long technical analysis can place different demands on the same machine. NVIDIA's published inference-cost guide makes workload and service-level assumptions part of the calculation. We would expect the same discipline in a private deployment proposal, with hardware, facilities and operating effort included. NVIDIA's inference-cost benchmarking guide, June 2025.
For some businesses, the difficulty is finding an owner for the ongoing work. Their IT team can specify access policies and approve infrastructure, but it may have little time to follow changes in serving engines or investigate why a model uses more memory after an update. Those tasks compete with work the team already has to deliver. Our service makes the inference operating work something the business can commission with an explicit scope and review over time.
That scope has to be precise about privacy. Keeping model processing inside the customer's infrastructure is one part of the arrangement. The application may still store conversations, retain uploaded files or call another service. Support permissions also need agreement. A useful design traces those paths and gives each one an owner. We define the inference responsibilities alongside the customer's application responsibilities so that neither disappears between suppliers.
For a business assessing a proposal, we suggest asking for five things in writing:
- The processing and storage locations, including those used by the application.
- The people and services authorised to access the environment, with the purpose of that access.
- The process for approving model changes and returning to a previous configuration.
- The workload and operating conditions against which performance will be assessed.
- The support and recovery responsibilities, including the access needed to carry them out.
LLM Machines' role is to put a supported inference configuration into that operating arrangement and keep it maintained. The customer retains control of its infrastructure and the decisions reserved to it. We provide workload discovery, sizing recommendations, configuration and ongoing support for the agreed inference environment, with a documented API for the applications the customer chooses.
The next article follows that work from the first workload samples through deployment, handover and maintenance: How LLM Machines delivers managed private inference.
